Merge remote-tracking branch 'build/master' into build-server

This commit is contained in:
igneum-labs 2026-10-08 11:24:02 +00:00
commit 1fa173b211
546 changed files with 104151 additions and 1873 deletions

View file

@ -5,7 +5,7 @@ tools: Read, Grep, Glob, Bash, Edit, Write, WebSearch, WebFetch, Agent
model: fable
---
You are the consensus engineer on a GPU-mined layer 1 built on a fork of rusty-kaspa. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
You are the consensus engineer on a GPU-mined layer 1 built on a fork of rusty-kaspa. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the founder changes them.
## What you carry in your head
The code of every major PoW node and how each one handles the problems you are about to meet:

View file

@ -5,7 +5,7 @@ tools: Read, Grep, Glob, Bash, WebSearch, WebFetch, Agent
model: fable
---
You are the cryptographer and proof-systems engineer on a GPU-mined layer 1 whose miners are also its ZK provers. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
You are the cryptographer and proof-systems engineer on a GPU-mined layer 1 whose miners are also its ZK provers. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the founder changes them.
## What you carry in your head
The whole history of proof of work and of proof systems, and you use it. When you make a claim about a chain, name the chain, the mechanism and where it lives in that chain's code. Examples of what you draw on:
@ -29,7 +29,7 @@ The whole history of proof of work and of proof systems, and you use it. When yo
- Numbers are measured or cited. A number from memory is labelled approximate. Never state an ASIC gain, a proving time or a verification time you have not measured or sourced.
- Write for an external reviewer: a spec section should let a stranger reproduce the argument.
- Prototype in Rust, with Metal on this Mac for GPU work and CUDA or OpenCL ports noted for miners. Benchmarks go in bench/ with the exact command and hardware.
- When you disagree with the design doc, say so once, with the attack or the measurement that drives it, then do the work under the doc's decision unless the project lead overrides.
- When you disagree with the design doc, say so once, with the attack or the measurement that drives it, then do the work under the doc's decision unless the founder overrides.
## Writing rules
No em dashes. Short sentences. Numbers in tables. The project is called Igneum. Approximate figures say so.

View file

@ -5,7 +5,7 @@ tools: Read, Grep, Glob, Bash, Edit, Write, WebSearch, WebFetch, Agent
model: fable
---
You are the execution engineer on a GPU-mined layer 1 whose every block is ZK-proven by its miners. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
You are the execution engineer on a GPU-mined layer 1 whose every block is ZK-proven by its miners. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the founder changes them.
## What you carry in your head
Every Ethereum client and every open zkVM, and what it costs to prove them:

View file

@ -5,7 +5,7 @@ tools: Read, Grep, Glob, Bash, WebSearch, WebFetch, Agent
model: fable
---
You are the miner-community lead on a GPU-mined layer 1 whose miners are also its ZK provers. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the project lead changes them.
You are the miner-community lead on a GPU-mined layer 1 whose miners are also its ZK provers. Read CLAUDE.md in the project root first, then the design doc it links to. The doc's decisions are fixed unless the founder changes them.
## What you carry in your head
Fifteen years of mining communities, and you remember what they did and why:

View file

@ -0,0 +1,26 @@
---
name: Grant application
about: Apply for an Igneum grant: tooling, a reference app, infrastructure or research, paid in IGN on delivery
title: "Grant: "
labels: grant
---
<!-- Read igneum.network/grants first. A short application is fine; the definition of done is the part that matters. -->
## What you will build
<!-- One paragraph. Name the tier: tooling, reference app, infrastructure, research. -->
## What done looks like
<!-- What runs where on the public testnet, and what a reviewer can click or call to see it working. -->
## Your public work
<!-- A repository, a package, a paper, a deployed thing. Links. -->
## How to reach you
<!-- An email, a Discord handle, or this issue. -->
<!-- Devnet and testnet coins have no value. No amount, price or date is promised. Not legal advice. -->

View file

@ -3,7 +3,7 @@
# ci.yml never enters it (7 October 2026: the inline `red` job of ci.yml was conditioned on master and release-*, and
# a feature branch would have waited for a merge of master before its reds were posted at all).
#
# One line per failed run (tools/ci/red-watch.mjs record, idempotent per run attempt) to /srv/ci-red/red.jsonl on the
# One line per failed, cancelled or timed-out run (tools/ci/red-watch.mjs record, idempotent per run attempt) to /srv/ci-red/red.jsonl on the
# box; the box's igneum-ci-red.timer posts each new line once to the hidden updates channel, naming the branch, the
# commit, the red check and the pushing author. Runs on the box's own runner (not a GitHub-hosted machine: the billing
# block of 6 October 2026, 18:37Z to 20:10Z, failed every hosted job at start and nobody was told). Never blocks a
@ -16,7 +16,9 @@ on:
jobs:
red:
name: red watcher (every branch; one line per failed run, with the branch, commit, red check and pushing author, to the updates channel and the box file)
if: ${{ github.event.workflow_run.conclusion == 'failure' }}
# failure, and since 7 October 2026 (17:2x UK) cancelled and timed_out too: a job that hangs into its timeout-minutes or a run
# someone cancels is a run that never answered, and a lane reads it like a red (tools/ci/red-watch.mjs names the kind)
if: ${{ github.event.workflow_run.conclusion == 'failure' || github.event.workflow_run.conclusion == 'cancelled' || github.event.workflow_run.conclusion == 'timed_out' }}
# the label ci-red is on igneum-build-1 only (added through the runners API on 7 October 2026; the default of
# RUNNER_LABELS in provision.sh carries it): the record file and the poster (igneum-ci-red.timer, the webhook file)
# live on that box, and the pool label igneum-build-1 is shared with igneum-build-2 since the same day
@ -35,6 +37,7 @@ jobs:
RED_WATCH_RUN_ID: ${{ github.event.workflow_run.id }}
RED_WATCH_ATTEMPT: ${{ github.event.workflow_run.run_attempt }}
RED_WATCH_WORKFLOW: ${{ github.event.workflow_run.name }}
RED_WATCH_CONCLUSION: ${{ github.event.workflow_run.conclusion }}
RED_WATCH_BRANCH: ${{ github.event.workflow_run.head_branch }}
RED_WATCH_SHA: ${{ github.event.workflow_run.head_sha }}
RED_WATCH_EVENT: ${{ github.event.workflow_run.event }}

View file

@ -36,6 +36,7 @@ jobs:
# code=true (no `before` to compare from), as does any error reading the compare API: when in doubt, run.
name: what the push touched (docs-only runs skip the Rust and simulator jobs)
runs-on: ubuntu-latest
timeout-minutes: 10 # a 7 s API call; every job carries a budget (tools/ci/workflow-timeouts-check.sh)
outputs:
code: ${{ steps.classify.outputs.code }}
steps:
@ -62,6 +63,7 @@ jobs:
needs: changes
if: ${{ needs.changes.outputs.code == 'true' }}
runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }}
timeout-minutes: 60 # the box's suite ran 45 s to 2 min 40 s on 7 October 2026; a hosted fallback compiles cold
steps:
- uses: actions/checkout@v4
- name: toolchain
@ -81,6 +83,7 @@ jobs:
# queue read 22); a feature-branch code push runs the igneum-pow tests alone. tools/ci/sims-branch-check.sh holds this rule.
if: ${{ needs.changes.outputs.code == 'true' && ((github.event_name == 'push' && (github.ref == 'refs/heads/master' || startsWith(github.ref, 'refs/heads/release-'))) || (github.event_name == 'pull_request' && (github.base_ref == 'master' || startsWith(github.base_ref, 'release-')))) }}
runs-on: ${{ vars.IGNEUM_CI_RUNNER == 'box' && fromJSON('["self-hosted", "linux", "x64", "igneum-build-1"]') || 'ubuntu-latest' }}
timeout-minutes: 45 # two simulators under 120 s each by their own timeout, plus a hosted fallback's pip install
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
@ -104,6 +107,10 @@ jobs:
site:
name: site build, link check, identity grep
runs-on: ubuntu-latest
# 15: the gate took 229 s on a hosted runner on 7 October 2026 plus a 40 s Playwright install; the same day three
# hosted site jobs on master hung in the gate for over two hours each with no budget, and GitHub's six-hour default
# would have ended each as a failure email. A hung job is a red the watcher posts (ci-red.yml fires on timed_out).
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
@ -118,4 +125,4 @@ jobs:
run: bash tools/ci/pre-push.sh --ci
- name: public stats API answers with the documented fields (the live site; master only, the endpoints exist there after the merge)
if: github.ref == 'refs/heads/master'
run: node tools/ci/public-api-check.mjs https://igneum.network
run: bash tools/ci/retry-once.sh public-api node tools/ci/public-api-check.mjs https://igneum.network # a live host: one retry before red

View file

@ -216,8 +216,14 @@ jobs:
run: |
$payload = Resolve-Path 'packaging\windows\igneum-windows-app'
function Run-Capture([string]$exe, [string]$flag) {
# Start-Process -Wait on an exe that exits in milliseconds can throw "Cannot process request because the process has
# exited" before it attaches (release-0.3.21 run 37653903394, 7 October 2026, 17:40 UK): one retry before the verdict.
$out = Join-Path $env:RUNNER_TEMP ('smoke-' + [IO.Path]::GetRandomFileName() + '.txt')
$p = Start-Process -FilePath $exe -ArgumentList $flag -Wait -NoNewWindow -PassThru -RedirectStandardOutput $out
$p = $null
foreach ($try in 1, 2) {
try { $p = Start-Process -FilePath $exe -ArgumentList $flag -Wait -NoNewWindow -PassThru -RedirectStandardOutput $out; break }
catch { if ($try -eq 2) { throw }; Write-Host ("Start-Process on {0} {1} failed once ({2}); second try" -f (Split-Path -Leaf $exe), $flag, $_.Exception.Message); Start-Sleep -Milliseconds 500 }
}
$text = if (Test-Path $out) { (Get-Content $out -Raw) } else { '' }
Write-Host ("{0} {1} -> exit {2}: {3}" -f (Split-Path -Leaf $exe), $flag, $p.ExitCode, $text.Trim())
if ($p.ExitCode -ne 0) { throw "$exe $flag exited $($p.ExitCode)" }

View file

@ -1,6 +1,6 @@
// Build script for igneum-app. On a Windows target it compiles resources/igneum-app.rc (the coin icon Explorer shows
// and the version block under Properties > Details) with windres and links the object into igneum-app.exe. Other
// targets: nothing. the project lead's rule (4 October 2026): every shipped exe carries the coin icon and a version block, like the
// targets: nothing. The founder's rule (4 October 2026): every shipped exe carries the coin icon and a version block, like the
// Mac app and DMG. No crate dependency: windres is called directly (x86_64-w64-mingw32-windres from Homebrew mingw-w64
// on the Mac, windres from MSYS2 on a PC; IGNEUM_WINDRES names another one).
use std::env;

View file

@ -1,6 +1,6 @@
// Windows resources for igneum-app.exe: the coin icon Explorer shows and the version block under Properties > Details.
// Compiled with x86_64-w64-mingw32-windres (the icon path is relative to brand/icons, passed with -I).
// the project lead's rule (4 October 2026): every shipped exe carries the coin icon and a version block, like the Mac app and DMG.
// the founder's rule (4 October 2026): every shipped exe carries the coin icon and a version block, like the Mac app and DMG.
#include <winver.h>
1 ICON "igneum.ico"

View file

@ -87,7 +87,7 @@ pub struct Settings {
/// (on when that is switched on, never effective while it is off). A pinned card is skipped.
#[serde(default)]
pub sweep: bool,
/// Power control (the project lead, 5 October 2026: "if we don't have to ask then don't ask"): the NVIDIA power cap and the
/// Power control (the founder, 5 October 2026: "if we don't have to ask then don't ask"): the NVIDIA power cap and the
/// efficiency sweep need administrator rights (one UAC prompt on Windows). Default OFF on every machine; the app
/// never raises the prompt on its own. Switching it on asks once, at that moment; a refused, cancelled or
/// unanswered prompt switches it back off with a notice, no retries.
@ -345,7 +345,7 @@ pub struct Runtime {
pub host: String,
/// Per-install random id (16 hex, app data dir/machine-id, locked to the user). Identity labels, vote keys and
/// the upload fields come from this, so two PCs cloned with the same COMPUTERNAME never share a key
/// (found 4 October 2026 on the project lead's two DESKTOP-KMCV30N machines).
/// (found 4 October 2026 on the founder's two DESKTOP-KMCV30N machines).
pub machine_id: String,
}
@ -357,7 +357,7 @@ impl Runtime {
let p2p_port = env("IGNEUM_APP_P2P_PORT").and_then(|v| v.parse().ok()).unwrap_or(26611);
let peers = match std::env::var("IGNEUM_APP_PEERS") {
Ok(v) => v.split(',').map(|s| s.trim().to_string()).filter(|s| !s.is_empty()).collect(),
// the public seed node first, then the project lead's Mac on the house LAN (devnet only)
// the public seed node first, then the founder's Mac on the house LAN (devnet only)
Err(_) if network == "devnet" => vec!["188.245.5.161:26611".to_string(), "192.168.68.64:26611".to_string()],
Err(_) => vec![],
};

View file

@ -23,8 +23,15 @@ pub const POWER_STEPS_PCT: [u32; 6] = [100, 90, 80, 70, 60, 50];
/// 6 October 2026, run 6: the 5090's best MH/W sat on the 60% floor (1,854 MHz: 0.563 MH/W, the rate within 0.15%),
/// so the ladder and the floor go to 45% of the maximum; the 1% rate tolerance is the guard below that
pub const CLOCK_STEPS_PCT: [u32; 7] = [100, 90, 80, 70, 60, 50, 45];
/// A card's clock floor when the vendor reports none: this share of its maximum core clock.
pub const CLOCK_FLOOR_PCT: u32 = 45;
/// A card's clock floor when the vendor reports none: this share of its maximum core clock. 7 October 2026, the PC 1
/// efficiency passes (docs/bench-log.md): the 5090's best MH per watt sat at 1,200 to 1,300 MHz (39 to 42 percent of
/// 3,090) and the rate fell past 5 percent only at 1,200 (class v3) and 1,100 (class v4), under the old 45 percent
/// floor; so the ladder continues below 45 percent in [`CLOCK_FINE_STEP_MHZ`] steps down to 20 percent of the maximum
/// (618 MHz on the 5090) and the stop rule, not the floor, ends the search.
pub const CLOCK_FLOOR_PCT: u32 = 20;
/// Below the percent ladder's last rung (45 percent) the clock ladder descends in steps of this many MHz until the
/// rate falls more than the tolerance under the cap point's rate (the knee), a step faults, or the floor is reached.
pub const CLOCK_FINE_STEP_MHZ: u32 = 100;
/// A point may lose this much rate against the fastest point and still win on MH per watt (the manifest can change it).
pub const RATE_TOLERANCE_PCT: f64 = 1.0;
/// A step whose hottest GPU reading reaches this is marked hot and cannot win (the engine aborts at 90).
@ -193,6 +200,9 @@ pub struct Plan {
pub tolerance_pct: f64,
power: Vec<Step>,
clock_pcts: Vec<u32>,
/// the clock ladder below the percent rungs: MHz values from the last rung minus one fine step down to the floor
/// (the core-clock knob of 7 October 2026; empty when the card has no readable maximum clock)
clock_fine: Vec<u32>,
fixed: Vec<Step>,
/// Ember 2 (Climb): the start point, the step sizes and the step budget
climb: Option<Climb>,
@ -218,7 +228,7 @@ impl Plan {
let mem_step = if limits.mem_max_mhz > limits.mem_default_mhz { ((limits.mem_max_mhz - limits.mem_default_mhz) / 20).max(25) } else { 0 };
let core_step = if limits.clock_max_mhz > 0 { (limits.clock_max_mhz / 20).max(25) } else { 0 };
let start = Point { clock_mhz: limits.clamp_clock(start.clock_mhz), power_pct: start.power_pct.clamp(50, 100), mem_mhz: limits.clamp_mem(start.mem_mhz) };
Plan { kind: PlanKind::Climb, limits: limits.clone(), before: start, tolerance_pct: goal.tolerance_pct(tolerance_pct), power: Vec::new(), clock_pcts: Vec::new(), fixed: Vec::new(), climb: Some(Climb { start, mem_step, core_step, budget: 5, goal }) }
Plan { kind: PlanKind::Climb, limits: limits.clone(), before: start, tolerance_pct: goal.tolerance_pct(tolerance_pct), power: Vec::new(), clock_pcts: Vec::new(), clock_fine: Vec::new(), fixed: Vec::new(), climb: Some(Climb { start, mem_step, core_step, budget: 5, goal }) }
}
/// The goal's score of a row: MH per watt for efficiency and balanced, the rate for maximum rate.
@ -285,7 +295,22 @@ impl Plan {
}
}
let clock_pcts = if limits.clock_max_mhz > 0 { CLOCK_STEPS_PCT[1..].to_vec() } else { Vec::new() };
Plan { kind: PlanKind::Full, limits: limits.clone(), before, tolerance_pct, power, clock_pcts, fixed: Vec::new(), climb: None }
// the fine ladder: from the last percent rung down to the floor in CLOCK_FINE_STEP_MHZ steps (the knob of
// 7 October 2026; the stop rule in `next` ends it at the knee)
let mut clock_fine = Vec::new();
if limits.clock_max_mhz > 0 {
let last_pct = limits.clamp_clock(limits.clock_max_mhz * CLOCK_STEPS_PCT[CLOCK_STEPS_PCT.len() - 1] / 100);
let floor = limits.clock_floor();
let mut m = (last_pct / CLOCK_FINE_STEP_MHZ) * CLOCK_FINE_STEP_MHZ;
if m >= last_pct {
m = m.saturating_sub(CLOCK_FINE_STEP_MHZ);
}
while m >= floor && m > 0 {
clock_fine.push(m);
m = m.saturating_sub(CLOCK_FINE_STEP_MHZ);
}
}
Plan { kind: PlanKind::Full, limits: limits.clone(), before, tolerance_pct, power, clock_pcts, clock_fine, fixed: Vec::new(), climb: None }
}
/// The prior's point, then one neighbour: the next clock step up when the prior caps the clock (is the cap
@ -304,14 +329,14 @@ impl Plan {
if neighbour != p {
fixed.push(Step { point: neighbour, watts: limits.watts_for(neighbour.power_pct), kind: Kind::Confirm });
}
Plan { kind: PlanKind::Confirm, limits: limits.clone(), before, tolerance_pct, power: Vec::new(), clock_pcts: Vec::new(), fixed, climb: None }
Plan { kind: PlanKind::Confirm, limits: limits.clone(), before, tolerance_pct, power: Vec::new(), clock_pcts: Vec::new(), clock_fine: Vec::new(), fixed, climb: None }
}
/// One step at the card's current point: the before number, and all a measure-only card (Apple, or NVIDIA
/// with Power control off) reports.
pub fn baseline(limits: &Limits, before: Point, tolerance_pct: f64) -> Plan {
let fixed = vec![Step { point: before, watts: limits.watts_for(before.power_pct), kind: Kind::Baseline }];
Plan { kind: PlanKind::Baseline, limits: limits.clone(), before, tolerance_pct, power: Vec::new(), clock_pcts: Vec::new(), fixed, climb: None }
Plan { kind: PlanKind::Baseline, limits: limits.clone(), before, tolerance_pct, power: Vec::new(), clock_pcts: Vec::new(), clock_fine: Vec::new(), fixed, climb: None }
}
/// How many steps the plan has at most (the clock ladder counts whether or not it runs).
@ -319,7 +344,42 @@ impl Plan {
if let Some(c) = &self.climb {
return c.budget;
}
self.fixed.len() + self.power.len() + self.clock_pcts.len()
self.fixed.len() + self.power.len() + self.clock_pcts.len() + self.clock_fine.len()
}
/// The cap point's row: the power ladder's choice (the row the clock search is read against), else the first row.
pub fn cap_row(&self, rows: &[Row]) -> Option<Row> {
if self.power.is_empty() {
rows.first().cloned()
} else {
choose(&rows[..self.power.len().min(rows.len())], self.tolerance_pct).or_else(|| rows.first().cloned())
}
}
/// Why the clock search ended after `rows`, in words, or None while it runs: a faulted clock row (a rejected or
/// mismatched hash during the hold: the fingerprint check), the knee (the rate under the cap point's by more than
/// the tolerance), or the floor.
pub fn clock_stop_reason(&self, rows: &[Row]) -> Option<String> {
let last = rows.last()?;
if last.point.clock_mhz == 0 || !matches!(self.kind, PlanKind::Full) {
return None;
}
let clock_rows = rows.len().saturating_sub(self.power.len());
if clock_rows == 0 {
return None;
}
if last.mark == Some(Mark::Faulted) {
return Some(format!("fingerprint mismatch at {} MHz, clocks reset", last.point.clock_mhz));
}
if let Some(cap) = self.cap_row(rows) {
if last.usable() && cap.usable() && cap.mhs > 0.0 && last.mhs < cap.mhs * (1.0 - self.tolerance_pct.max(0.0) / 100.0) {
return Some(format!("rate fell {:.1} percent at {} MHz", 100.0 * (cap.mhs - last.mhs) / cap.mhs, last.point.clock_mhz));
}
}
if clock_rows >= self.clock_pcts.len() + self.clock_fine.len() {
return Some(format!("the floor at {} MHz", last.point.clock_mhz));
}
None
}
pub fn is_empty(&self) -> bool {
self.len() == 0
@ -338,10 +398,22 @@ impl Plan {
return Some(self.power[i].clone());
}
let k = i - self.power.len();
let pct = *self.clock_pcts.get(k)?;
// the stop rule (7 October 2026): a faulted clock row or the knee ends the search; the choice is made among
// the rows so far
if k > 0 {
if let Some(reason) = self.clock_stop_reason(rows) {
if !reason.starts_with("the floor") {
return None;
}
}
}
let clock = if k < self.clock_pcts.len() {
self.limits.clamp_clock(self.limits.clock_max_mhz * self.clock_pcts[k] / 100)
} else {
*self.clock_fine.get(k - self.clock_pcts.len())?
};
// the clock ladder rides the power point the power ladder chose (the before point when nothing won)
let power_pct = if self.power.is_empty() { self.before.power_pct } else { choose(&rows[..self.power.len()], self.tolerance_pct).map(|r| r.point.power_pct).unwrap_or(self.before.power_pct) };
let clock = self.limits.clamp_clock(self.limits.clock_max_mhz * pct / 100);
// a step whose clamp lands on the previous step's clock is dropped (the floor was reached)
if rows.last().map(|r| r.point.clock_mhz == clock).unwrap_or(false) {
return None;
@ -921,6 +993,38 @@ pub fn result_line(kind: PlanKind, mhs: f64, watts: f64, eff: f64) -> String {
}
}
/// The core-clock knob's result on a card (7 October 2026; the UI lane's field shape): the chosen lock against the cap
/// point's unlocked row, and the stop reason in words.
#[derive(Clone, Debug, Default, PartialEq)]
pub struct LockResult {
/// the chosen core clock cap (0 = unlocked)
pub lock_mhz: u32,
pub lock_mhs: f64,
pub lock_w: f64,
pub lock_mhw: f64,
/// the cap point's row (clock 0): the rate and draw the lock is read against
pub unlocked_mhs: f64,
pub unlocked_w: f64,
/// "rate fell 5.1 percent at 1,200 MHz", "fingerprint mismatch at 1,400 MHz, clocks reset", "the floor at 618 MHz",
/// "no lever" (a card without a clock cap), "" while nothing ran
pub lock_note: String,
}
/// The knob's result from a finished plan's rows and its chosen row.
pub fn lock_result(plan: &Plan, rows: &[Row], chosen: &Row) -> LockResult {
let cap = plan.cap_row(rows);
let (unlocked_mhs, unlocked_w) = cap.as_ref().map(|c| (c.mhs, c.watts)).unwrap_or((0.0, 0.0));
let ran_clocks = rows.iter().any(|r| r.point.clock_mhz > 0);
let note = if plan.limits.clock_max_mhz == 0 {
"no lever".to_string()
} else if !ran_clocks {
String::new()
} else {
plan.clock_stop_reason(rows).unwrap_or_else(|| format!("stopped at {} MHz", rows.last().map(|r| r.point.clock_mhz).unwrap_or(0)))
};
LockResult { lock_mhz: chosen.point.clock_mhz, lock_mhs: chosen.mhs, lock_w: chosen.watts, lock_mhw: chosen.eff, unlocked_mhs, unlocked_w, lock_note: note }
}
/// Why a card cannot be tuned beyond measuring, or None when both knobs are available.
pub fn control_reason(vendor: &str, limits: &Limits, device: &str, power_control: bool, amd_helper: bool) -> Option<String> {
match vendor {
@ -952,7 +1056,9 @@ mod tests {
#[test]
fn the_full_plan_is_the_power_ladder_then_the_clock_ladder_at_the_chosen_power() {
let plan = Plan::full(&l5090(), Point { clock_mhz: 0, power_pct: 80, mem_mhz: 0 }, 1.0);
assert_eq!(plan.len(), 5 + 6, "five power steps (60% and 50% clamp to 400 W; one kept) and six clock steps (90% down to 45%)");
// five power steps (60% and 50% clamp to 400 W; one kept), six percent rungs (90% down to 45% = 1,390) and the
// fine ladder 1,300 down to the 20% floor (618): 1,300, 1,200, ..., 700 = 7 steps
assert_eq!(plan.len(), 5 + 6 + 7);
let first = plan.next(&[]).unwrap();
assert_eq!((first.point, first.watts, first.kind), (Point { clock_mhz: 0, power_pct: 100, mem_mhz: 0 }, 575.0, Kind::Power));
// the power ladder: 575, 518, 460, 403, 400
@ -974,24 +1080,21 @@ mod tests {
let s = plan.next(&rows).unwrap();
assert_eq!(s.point.clock_mhz, 2472);
rows.push(row_at(s.point, 220.0, 123.5));
rows.push(row_at(plan.next(&rows).unwrap().point, 200.0, 118.0));
let s = plan.next(&rows).unwrap();
assert_eq!(s.point.clock_mhz, 1854, "60% of 3,090");
rows.push(row_at(s.point, 180.0, 100.0));
let s = plan.next(&rows).unwrap();
assert_eq!(s.point.clock_mhz, 1545, "50%");
rows.push(row_at(s.point, 170.0, 90.0));
let s = plan.next(&rows).unwrap();
assert_eq!(s.point.clock_mhz, 1390, "45% of 3,090 is the floor (6 October 2026)");
rows.push(row_at(s.point, 160.0, 80.0));
assert_eq!(s.point.clock_mhz, 2163, "70%");
rows.push(row_at(s.point, 200.0, 118.0));
// the stop rule (7 October 2026): 118 is 4.8% under the cap point's 124, past the 1% tolerance, so the search
// ends here (the 6 October ladder went on to 1,854, 1,545 and 1,390)
assert_eq!(plan.next(&rows), None);
assert_eq!(plan.clock_stop_reason(&rows).as_deref(), Some("rate fell 4.8 percent at 2163 MHz"));
// the choice: 2,472 MHz keeps 99.6% of the top rate at 220 W = 0.561 MH/W; 2,163 MHz (118 MH/s) is outside the 1% tolerance
let best = choose(&rows, 1.0).unwrap();
assert_eq!(best.point, Point { clock_mhz: 2472, power_pct: 100, mem_mhz: 0 });
// a wider tolerance lets the 2,163 MHz step (0.590 MH/W, 4.8% slower) win
assert_eq!(choose(&rows, 5.0).unwrap().point.clock_mhz, 2163);
// no power limits, clocks only; no clocks, power only; nothing, empty
assert_eq!(Plan::full(&Limits { clock_max_mhz: 2000, ..Default::default() }, Point::default(), 1.0).len(), 6);
// clocks only: six percent rungs (1,800 .. 900) then the fine ladder 800 .. 400 (the 20% floor) = 5 more
assert_eq!(Plan::full(&Limits { clock_max_mhz: 2000, ..Default::default() }, Point::default(), 1.0).len(), 6 + 5);
assert_eq!(Plan::full(&Limits { power_default_w: 300.0, ..Default::default() }, Point::default(), 1.0).len(), 6);
assert!(Plan::full(&Limits::default(), Point::default(), 1.0).is_empty());
}
@ -1017,11 +1120,12 @@ mod tests {
#[test]
fn limits_never_exceed_the_vendor_or_undercut_the_floor() {
let l = l5090();
assert_eq!(l.clock_floor(), 1390);
assert_eq!(l.clamp_clock(1000), 1390);
assert_eq!(l.clock_floor(), 618, "20% of 3,090 (7 October 2026; the 45% floor of 6 October sat on the 5090's knee)");
assert_eq!(l.clamp_clock(1000), 1000);
assert_eq!(l.clamp_clock(500), 618);
assert_eq!(l.clamp_clock(5000), 3090);
assert_eq!(l.clamp_clock(0), 0, "unlocked stays unlocked");
assert_eq!(Limits { clock_max_mhz: 3000, clock_min_mhz: 2100, ..Default::default() }.clamp_clock(1500), 2100, "the vendor's floor wins over the 45% rule");
assert_eq!(Limits { clock_max_mhz: 3000, clock_min_mhz: 2100, ..Default::default() }.clamp_clock(1500), 2100, "the vendor's floor wins over the 20% rule");
assert_eq!(l.watts_for(100), 575.0);
assert_eq!(l.watts_for(50), 400.0);
assert_eq!(Limits { power_default_w: 300.0, power_max_w: 250.0, ..Default::default() }.watts_for(100), 250.0);
@ -1321,4 +1425,138 @@ mod tests {
assert!(control_reason("amd", &Limits::default(), "1", false, true).is_none());
}
}
/// The core-clock knob (7 October 2026, the PC 1 efficiency passes): a flat ladder walks below the old 45 percent
/// floor in 100 MHz steps to the 20 percent floor, and the result names the floor.
#[test]
fn the_clock_ladder_continues_below_45_percent_in_100_mhz_steps_to_the_floor() {
let plan = Plan::full(&l5090(), Point { clock_mhz: 0, power_pct: 100, mem_mhz: 0 }, 1.0);
let mut rows = Vec::new();
for _ in 0..5 {
let s = plan.next(&rows).unwrap();
rows.push(row_at(s.point, 300.0, 136.8));
}
let mut clocks = Vec::new();
while let Some(s) = plan.next(&rows) {
assert_eq!(s.kind, Kind::Clock);
clocks.push(s.point.clock_mhz);
// the rate holds (memory-bound): the draw falls with the clock
rows.push(row_at(s.point, 300.0 - clocks.len() as f64 * 10.0, 136.0));
}
assert_eq!(clocks, vec![2781, 2472, 2163, 1854, 1545, 1390, 1300, 1200, 1100, 1000, 900, 800, 700]);
assert_eq!(plan.clock_stop_reason(&rows).as_deref(), Some("the floor at 700 MHz"));
let chosen = choose(&rows, 1.0).unwrap();
assert_eq!(chosen.point.clock_mhz, 700, "flat rate: the lowest draw wins");
let r = lock_result(&plan, &rows, &chosen);
assert_eq!((r.lock_mhz, r.lock_w, r.unlocked_mhs, r.unlocked_w), (700, 170.0, 136.8, 300.0));
assert_eq!(r.lock_note, "the floor at 700 MHz");
}
/// The stop rule on PC 1's RTX 5090 rows of 7 October 2026 (class v3, the card alone): the first clock row more
/// than the tolerance under the cap point's rate ends the search and the best MH per watt among the rows within
/// tolerance is chosen. At the 1 percent tolerance the 5090's rate (136.6 at 2,781) is 1.24 percent down at
/// 1,545 MHz, so the search ends there and 1,854 MHz (135.6 MH/s at 239.6 W) is the point; at 1.5 percent it runs
/// on to 1,200 (5.2 percent down) and 1,300 MHz is the point. The tolerance is the manifest's.
#[test]
fn the_clock_search_stops_at_the_knee_and_names_it() {
let measured: Vec<(u32, f64, f64)> = vec![(2781, 317.9, 136.6), (2472, 276.0, 136.4), (2163, 252.4, 136.1), (1854, 239.6, 135.6), (1545, 232.1, 134.9), (1390, 229.0, 134.85), (1300, 223.3, 134.6), (1200, 215.7, 129.5)];
let walk = |tolerance: f64| -> (Plan, Vec<Row>) {
let plan = Plan::full(&Limits { clock_max_mhz: 3090, ..Default::default() }, Point { clock_mhz: 0, power_pct: 100, mem_mhz: 0 }, tolerance);
let mut rows = vec![];
for (mhz, w, mhs) in &measured {
let Some(s) = plan.next(&rows) else { break };
assert_eq!(s.point.clock_mhz, *mhz);
rows.push(row_at(s.point, *w, *mhs));
}
(plan, rows)
};
let (plan, rows) = walk(1.0);
assert_eq!(rows.last().unwrap().point.clock_mhz, 1545, "the search ends on the first row over 1 percent under the cap row");
assert_eq!(plan.next(&rows), None);
let reason = plan.clock_stop_reason(&rows).unwrap();
assert!(reason.starts_with("rate fell 1.2 percent at 1545 MHz"), "{reason}");
let chosen = choose(&rows, 1.0).unwrap();
assert_eq!(chosen.point.clock_mhz, 1854, "the best MH per watt within 1 percent of the fastest row");
let r = lock_result(&plan, &rows, &chosen);
assert_eq!((r.lock_mhz, r.unlocked_mhs), (1854, 136.6));
assert!((r.lock_mhw - 135.6 / 239.6).abs() < 1e-6);
let (plan, rows) = walk(1.5);
assert_eq!(rows.last().unwrap().point.clock_mhz, 1200);
assert_eq!(plan.next(&rows), None);
assert!(plan.clock_stop_reason(&rows).unwrap().starts_with("rate fell 5.2 percent at 1200 MHz"));
assert_eq!(choose(&rows, 1.5).unwrap().point.clock_mhz, 1300);
}
/// The fingerprint rule: a clock row marked Faulted (a rejected or mismatched hash during the hold) ends the search
/// at once; the choice is made among the usable rows and the note says why.
#[test]
fn a_faulted_clock_row_ends_the_search_and_the_note_says_so() {
let plan = Plan::full(&Limits { clock_max_mhz: 3090, ..Default::default() }, Point { clock_mhz: 0, power_pct: 100, mem_mhz: 0 }, 1.0);
let mut rows = vec![];
for (mhz, w, mhs) in [(2781, 317.9, 136.6), (2472, 276.0, 136.4)] {
let s = plan.next(&rows).unwrap();
assert_eq!(s.point.clock_mhz, mhz);
rows.push(row_at(s.point, w, mhs));
}
let s = plan.next(&rows).unwrap();
assert_eq!(s.point.clock_mhz, 2163);
let mut bad = row_at(s.point, 252.4, 136.1);
bad.faults = 1;
bad.mark = Some(Mark::Faulted);
rows.push(bad);
assert_eq!(plan.next(&rows), None, "the search ends on the faulted row");
assert_eq!(plan.clock_stop_reason(&rows).as_deref(), Some("fingerprint mismatch at 2163 MHz, clocks reset"));
let chosen = choose(&rows, 1.0).unwrap();
assert_eq!(chosen.point.clock_mhz, 2472, "the faulted row never wins");
assert_eq!(lock_result(&plan, &rows, &chosen).lock_note, "fingerprint mismatch at 2163 MHz, clocks reset");
}
/// The same through the state machine with a fake helper (the known-failed case first: a mismatch mid-search must
/// reset and abort): the run applies 2,781 and 2,472, a fault lands during 2,163's hold, the row comes out Faulted,
/// the next step is none, and the run's final Apply is the chosen 2,472 point (the reset), then Finished.
#[test]
fn a_mismatch_mid_search_resets_to_the_chosen_point_and_finishes() {
let plan = Plan::full(&Limits { clock_max_mhz: 3090, ..Default::default() }, Point { clock_mhz: 0, power_pct: 100, mem_mhz: 0 }, 1.0);
let timing = Timing { settle: Duration::from_secs(1), hold: Duration::from_secs(2), apply: Duration::from_secs(3) };
let t0 = Instant::now();
let mut run = Run::new(0, "0", "card-0", plan, 300.0, false, timing, t0);
let mut t = t0;
let mut applied: Vec<Step> = Vec::new();
let mut finished: Option<Row> = None;
let watts_for = |mhz: u32| -> f64 { match mhz { 2781 => 317.9, 2472 => 276.0, _ => 252.4 } };
for _ in 0..200 {
t += Duration::from_millis(500);
let acked = true;
let limit = run.current.as_ref().map(|s| s.watts).unwrap_or(0.0);
// the fake helper: every setting takes; during 2,163's hold the worker reports a mismatched hash
if let Some(cur) = run.current.clone() {
if matches!(run.phase, Phase::Holding { .. }) {
run.sample_rate(136.4);
run.sample_telemetry(watts_for(cur.point.clock_mhz), cur.point.clock_mhz as f64, 13801.0, 60.0);
if cur.point.clock_mhz == 2163 {
run.sample_fault();
}
}
}
for o in run.tick(t, Readback { limit_w: limit, acked }) {
match o {
Out::Apply(s) => applied.push(s),
Out::Finished(r) => finished = Some(r),
Out::Failed(e) => panic!("the run failed: {e}"),
Out::Row(_) => {}
}
}
if finished.is_some() {
break;
}
}
let clocks: Vec<u32> = applied.iter().map(|s| s.point.clock_mhz).collect();
assert_eq!(clocks, vec![2781, 2472, 2163, 2472], "2,781, 2,472, the faulted 2,163, then the reset to the chosen 2,472");
assert_eq!(applied.last().unwrap().kind, Kind::Confirm);
let f = finished.expect("finished");
assert_eq!(f.point.clock_mhz, 2472);
assert_eq!(run.rows.len(), 3);
assert_eq!(run.rows[2].mark, Some(Mark::Faulted));
assert_eq!(run.plan.clock_stop_reason(&run.rows).as_deref(), Some("fingerprint mismatch at 2163 MHz, clocks reset"));
}
}

View file

@ -2727,6 +2727,13 @@ impl Engine {
cc.tune_steps = of;
cc.tune_eta_s = eta;
cc.tune_plan = plan.into();
// the clock search's own step count while a clock step runs (the UI's "locking clocks: step 4 of 9")
if let Some(r) = self.sweep.as_ref() {
let power_steps = r.rows.iter().filter(|x| x.point.clock_mhz == 0 && x.mark.is_some()).count() as u32;
let on_clock = r.current.as_ref().map(|s| s.kind == crate::ember::Kind::Clock).unwrap_or(false);
cc.lock_step = if on_clock { step.saturating_sub(power_steps) } else { 0 };
cc.lock_steps = if on_clock { of.saturating_sub(power_steps) } else { 0 };
}
}
if self.shared.runtime.sweep_only {
// the job playbook forwards this to the installed app's /api/tune-progress
@ -2765,6 +2772,22 @@ impl Engine {
c.tune_source = kind.name().into();
c.tune_line = crate::ember::result_line(kind, row.mhs, row.watts, row.eff);
c.tune_curve = run.rows.iter().map(|r| r.json()).collect();
// the core-clock knob's result (7 October 2026): the chosen lock against the cap point, and why it stopped
if kind == crate::ember::PlanKind::Baseline {
c.lock_note = if c.vendor == "apple" || run.plan.limits.clock_max_mhz == 0 { "no lever".into() } else { c.sweep_note.clone() };
} else {
let lr = crate::ember::lock_result(&run.plan, &run.rows, &row);
c.lock_mhz = lr.lock_mhz;
c.lock_mhs = lr.lock_mhs;
c.lock_w = lr.lock_w;
c.lock_mhw = lr.lock_mhw;
c.unlocked_mhs = lr.unlocked_mhs;
c.unlocked_w = lr.unlocked_w;
c.lock_at = unix as f64;
c.lock_note = lr.lock_note;
}
c.lock_step = 0;
c.lock_steps = 0;
let control = c.tune_control;
if kind == crate::ember::PlanKind::Baseline {
c.sweep_note = if control { String::new() } else { c.sweep_note.clone() };
@ -4079,7 +4102,7 @@ impl Engine {
}
/// The watts a card's cap asks for: power_pct of the default limit, inside the card's min and max.
/// the project lead, 5 October 2026: "if we don't have to ask then don't ask". The NVIDIA power cap and the efficiency sweep need
/// the founder, 5 October 2026: "if we don't have to ask then don't ask". The NVIDIA power cap and the efficiency sweep need
/// administrator rights (one UAC prompt on Windows, pkexec on Linux); the engine builds an elevated command only when
/// Power control is on in Settings, or when it is itself the elevated PC sweep job (--sweep).
fn elevation_allowed(power_control: bool, sweep_only: bool) -> bool {
@ -4351,7 +4374,7 @@ mod resume_tests {
mod tests {
#[test]
fn power_control_off_builds_no_elevated_command() {
// the decision (the project lead, 5 October 2026): off = the app never asks; the elevated PC sweep job is the exception
// the decision (the founder, 5 October 2026): off = the app never asks; the elevated PC sweep job is the exception
assert!(!super::elevation_allowed(false, false));
assert!(super::elevation_allowed(true, false));
assert!(!super::elevation_allowed(false, true), "the --sweep job alone never asks (C35)");

View file

@ -1,5 +1,5 @@
//! The `build` job (the model is src/jobs.rs, the runner src/jobrun.rs): a Windows PC builds the node and the app
//! engine for Linux and Windows inside its WSL2 Ubuntu, as root, with nothing from the project lead. the project lead's ask, 4 October 2026
//! engine for Linux and Windows inside its WSL2 Ubuntu, as root, with nothing from the founder. The founder's ask, 4 October 2026
//! evening ("efficiency"): every Windows build went through a GitHub runner at 15 to 25 minutes a round and every
//! Linux binary was cross-compiled on the Mac under the build lock; the two RTX 5090 PCs sit idle on the CPU side.
//!

View file

@ -32,7 +32,7 @@ use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant};
/// The safety-net poll. Before 0.3.6 this was 600 s and a published job waited up to 10 minutes on every PC (the project lead,
/// The safety-net poll. Before 0.3.6 this was 600 s and a published job waited up to 10 minutes on every PC (the founder,
/// 5 October 2026: "why is it taking so long for pc2 and pc1s tasks to spin up?"); the wake below makes it seconds.
const CHECK_EVERY_S: u64 = 120;
const RETRY_AFTER_ERROR_S: u64 = 300;
@ -50,7 +50,7 @@ const GPU_IDLE_PCT: f64 = 5.0;
const GPU_IDLE_WAIT_S: u64 = 180;
const DEFAULT_DISTRO: &str = "Ubuntu-24.04";
/// WSL jobs run as root by default: the Ubuntu the app sees is the one of the account the app runs under, and a
/// personal user ([user] on PC 2) need not exist there (4 October 2026: `getpwnam([user]) failed`).
/// personal user (<user> on PC 2) need not exist there (4 October 2026: `getpwnam(<user>) failed`).
const DEFAULT_WSL_USER: &str = "root";
const DEFAULT_FIXTURES: &[&str] = &["block-338-shard1", "block-341-shards2", "block-344-shards4"];
const HISTORY_SHOWN: usize = 20;

View file

@ -2,7 +2,7 @@
//! manifest on the downloads host, and every Igneum Miner app polls it (src/jobrun.rs, every 10 minutes). A job
//! runs at most once per id on a machine, only when its target matches (machine id, platform, requirements) and
//! it has not expired. Same key, same canonical JSON (sorted keys, no whitespace) and the same `.sig` scheme as the
//! update manifest (src/manifest.rs). the project lead's rule, 4 October 2026: one app on both PCs that the Mac can send
//! update manifest (src/manifest.rs). The founder's rule, 4 October 2026: one app on both PCs that the Mac can send
//! commands and files to over the line, so everything is tested and built without a person at the PC.
//!
//! This module is self-contained (serde_json and manifest.rs only), so the signer (src/bin/ota-sign.rs) includes it

View file

@ -1,4 +1,4 @@
//! Over-the-air updates of the app (and the node, miner and workers inside it). the project lead's rule: every app updates
//! Over-the-air updates of the app (and the node, miner and workers inside it). The founder's rule: every app updates
//! itself and downloads the update without being asked. This is also how a consensus upgrade (a height-activated
//! rule such as difficulty v2) reaches every node before its activation height.
//!
@ -120,7 +120,7 @@ pub struct Updater {
/// this machine's minute of the hour for applying (manifest::slot_minute of the machine id)
slot: u64,
/// When this engine started (unix seconds): an update published more than an hour before it is a catch-up, not a
/// rollout, and skips the hourly slot (the project lead's morning of 6 October 2026: PC 1 came up after the 0.3.11 publish and
/// rollout, and skips the hourly slot (the founder's morning of 6 October 2026: PC 1 came up after the 0.3.11 publish and
/// sat on "installs at the next safe moment" until he pressed Install now).
started_unix: u64,
catch_up_logged: bool,

View file

@ -1,4 +1,4 @@
//! One administrator approval, ever (the project lead, 6 October 2026, 11:50 UTC, after clicking the third prompt of the morning:
//! One administrator approval, ever (the founder, 6 October 2026, 11:50 UTC, after clicking the third prompt of the morning:
//! "can we make sure all these popups are not needed in future?").
//!
//! What 0.3.12 does: Power control on raises one prompt and sets every cap in that step; but every later cap (an app

View file

@ -1,4 +1,4 @@
//! Proving v1 step 1 (5 October 2026, the project lead: "open the proving round asap"): the prover is on by default on every
//! Proving v1 step 1 (5 October 2026, the founder: "open the proving round asap"): the prover is on by default on every
//! mining machine that can prove, decided once per install after the cards are detected (src/engine.rs
//! `apply_prove_default`). The rule, one line each:
//!
@ -6,13 +6,13 @@
//! |---|---|---|
//! | NVIDIA card with 24 GB or more, mining or not, Windows with WSL2 (Ubuntu-24.04) answering or Linux | on | a full shard at the adopted v1 budget (30,000 pgas, 4.7 M cycles) peaks at 20,434 MiB alone and 22,210 beside the miner (measured on the 5090; approximate for a 24 GB card's own allocation); the prototype shard the devnet proves until its fee switch (6.75 M pgas) peaks at 28,307 MiB alone and 30,039 beside the miner, so until the switch only a 32 GB card proves it and a 24 GB card's prover waits for shards it can hold (the host refuses nothing; a proof that runs out of memory fails and the shard is left) |
//! | NVIDIA card of 16 to 24 GB | off, with the line saying why | the GPU prover's floor is 13,874 MiB for an EMPTY shard, 15,670 beside the miner; a 16 GB card holds no full shard |
//! | NVIDIA card under 16 GB | off | 13,874 MiB does not fit; the project lead's 12 GB requirement is open until a prover build with a smaller floor is measured |
//! | NVIDIA card under 16 GB | off | 13,874 MiB does not fit; the founder's 12 GB requirement is open until a prover build with a smaller floor is measured |
//! | Windows under 32 GB of RAM | off, with the line saying why | the WSL2 prover held 7.9 GB on a 63 GB PC; a 16 GB PC would swap |
//! | Windows with a qualifying card but WSL2 silent | off, with the Set up hint | nothing can prove until the distribution exists |
//! | Apple silicon | off | the M5 Max CPU took 41 to 55 s for an EMPTY shard's compressed proof under load and 272 s for a 200-pgas shard; a full shard was never under 60 s (bench-log 4 and 5 October 2026) |
//! | AMD-only (no NVIDIA card) | off, "mines and does not prove" | no zkVM proves on an AMD GPU today (docs/analysis/amd-proving.md); the SP1 CPU prover on PC 1 cost 82 to 87 s core plus 199 to 202 s compressed a shard at a 30 GB RSS whatever the shard size (bench-log, "the SP1 CPU prover on PC 1") |
//!
//! Decided 5 October 2026 (delegated by the project lead: "deploy what is absolute best"), docs/plans/proving-v1.md. The default
//! Decided 5 October 2026 (delegated by the founder: "deploy what is absolute best"), docs/plans/proving-v1.md. The default
//! never switches an explicit on back off, and Settings always wins afterwards.
use crate::state::CardState;

View file

@ -127,7 +127,7 @@ pub struct CardState {
pub tune_source: String, // full | confirm | baseline
pub tune_line: String, // "Tuned: 122.3 MH/s at 290 W (0.422 MH/W)" once tuned
// a tune in progress on this card, by this engine or by a measurement engine posting /api/tune-progress
// (the project lead, 6 October 2026: "don't we need to show in the app that tuning is in progress?")
// (the founder, 6 October 2026: "don't we need to show in the app that tuning is in progress?")
pub tune_step: u32,
pub tune_steps: u32,
pub tune_eta_s: i64,
@ -135,6 +135,18 @@ pub struct CardState {
// Ember 2: the memory clock the last tune chose and the measured curve (every row of the last plan)
pub tune_mem_mhz: u32,
pub tune_curve: Vec<serde_json::Value>,
// the core-clock knob (7 October 2026, src/ember.rs lock_result; the UI lane's field shape): the chosen lock
// against the cap point's unlocked row, the step while the clock search runs, the moment and the stop reason
pub lock_mhz: u32, // the chosen core clock cap (0 = unlocked)
pub lock_mhs: f64,
pub lock_w: f64,
pub lock_mhw: f64,
pub unlocked_mhs: f64, // the cap point's row the lock is read against
pub unlocked_w: f64,
pub lock_step: u32, // the clock search's step while it runs (0 otherwise)
pub lock_steps: u32,
pub lock_at: f64, // unix s the lock point was taken (0 = never)
pub lock_note: String, // "rate fell 5.1 percent at 1200 MHz", "fingerprint mismatch at 1400 MHz, clocks reset", "the floor at 700 MHz", "no lever"
// the kernel variant race (docs/design/miner-tuning.md): what the worker's last race chose
pub variant: String,
pub race_mhs: f64,

View file

@ -221,11 +221,11 @@ mod tests {
#[test]
fn the_command_line_after_the_dashes_never_carries_a_double_quote_or_a_newline() {
let file = Path::new("C:\\Users\\the project lead\\AppData\\Local\\igneum\\wsl\\probe-12-3.sh");
let line = bash_line(file, true, &["--proof", "/mnt/c/Users/the project lead/AppData/Local/igneum/app/proving/p.bin", "--statement", "0xab", "it's", "two\nlines"]);
let file = Path::new("C:\\Users\\the founder\\AppData\\Local\\igneum\\wsl\\probe-12-3.sh");
let line = bash_line(file, true, &["--proof", "/mnt/c/Users/the founder/AppData/Local/igneum/app/proving/p.bin", "--statement", "0xab", "it's", "two\nlines"]);
assert_eq!(
line,
"bash -l '/mnt/c/Users/the project lead/AppData/Local/igneum/wsl/probe-12-3.sh' '--proof' '/mnt/c/Users/the project lead/AppData/Local/igneum/app/proving/p.bin' '--statement' '0xab' 'it'\\''s' 'two lines'"
"bash -l '/mnt/c/Users/the founder/AppData/Local/igneum/wsl/probe-12-3.sh' '--proof' '/mnt/c/Users/the founder/AppData/Local/igneum/app/proving/p.bin' '--statement' '0xab' 'it'\\''s' 'two lines'"
);
assert!(!line.contains('"') && !line.contains('\n'), "{line}");
// the lookup script's own double quotes live in the file, never on the line
@ -282,7 +282,7 @@ mod tests {
#[test]
fn wsl_paths() {
assert_eq!(wsl_path(Path::new("C:\\Users\\[user]\\AppData\\Local\\igneum\\app\\proving\\seq.json")), "/mnt/c/Users/[user]/AppData/Local/igneum/app/proving/seq.json");
assert_eq!(wsl_path(Path::new("C:\\Users\\<user>\\AppData\\Local\\igneum\\app\\proving\\seq.json")), "/mnt/c/Users/<user>/AppData/Local/igneum/app/proving/seq.json");
assert_eq!(wsl_path(Path::new("\\\\?\\D:\\x")), "/mnt/d/x");
assert_eq!(wsl_path(Path::new("/tmp/x")), "/tmp/x");
}

View file

@ -267,7 +267,7 @@ var View = (function () {
function cap(t) { return t ? t.charAt(0).toUpperCase() + t.slice(1) : ''; }
var kindWord = Notices.kindWord;
// hot-plug (src/hotplug.rs): a removed card's row hides after five minutes (gone); a faulty one has no switch
// The list is ordered by performance (the project lead, 6 October 2026): usable cards first, then by the measured rate since the
// The list is ordered by performance (the founder, 6 October 2026): usable cards first, then by the measured rate since the
// start (5 MH/s buckets so the order does not flicker), then discrete, external and Apple before integrated, then
// memory; removed and unusable cards last. Ties keep the detection order.
function perfRank(c) {

View file

@ -1,6 +1,6 @@
# Igneum brand assets
## The rule (the project lead, 4 October 2026)
## The rule (the founder, 4 October 2026)
One global logo for apps, profile pictures, favicons, everything. The mark sits in a black square. Never in a circle.
Never on another colour. Minimum clear space = 20% of the square. The Mac app icon is the model: a full square, all

View file

@ -1,7 +1,7 @@
#!/usr/bin/env python3
"""Builds every Igneum icon from the master mark, brand/master/igneum-mark-square.svg (4 October 2026).
the project lead's rule (4 October 2026): one global logo for apps, profile pictures, favicons, everything. The mark sits in a black
The founder's rule (4 October 2026): one global logo for apps, profile pictures, favicons, everything. The mark sits in a black
square (#0C0C0E, the site's obsidian token), centred, no circle, no ring, no border, never on another colour. The Mac app
is the model. The rounded variant (brand/master/igneum-mark-square-rounded.svg, Apple's 824-on-1024 icon grid) is used
only where the shape has to be baked into the file: the DMG volume icon. Tahoe masks the plain square itself.

View file

@ -15,7 +15,7 @@ Mark description for the device, when a form asks: "A stylised flame formed of t
## Applicant
Igneum Labs LTD, Licensee Address: Unit IH-00-01-01-OF-01, Level 01, Innovation One, Dubai International Financial
Centre (decided 5 October 2026; it replaces the ADGM DLT Foundation named on 4 October). NOT [other-business]: a [other-business]
Centre (decided 5 October 2026; it replaces the ADGM DLT Foundation named on 4 October). NOT the earlier entity: an earlier-entity
filing would tie Igneum to VIVA and to its owner through public registers, which the standing rule forbids. The
registered address above is the applicant address, with a trademark attorney as the address for service, so no
personal address appears anywhere. If the company's registration is not complete the week you want to file, the

View file

@ -0,0 +1,18 @@
[profile.default]
src = "src"
test = "test"
script = "script"
out = "out"
libs = []
solc_version = "0.8.28"
# The verifier calls the BLS12-381 precompiles of EIP-2537 (live on Sepolia and mainnet since Pectra), so the test EVM
# runs the Osaka rules, the fork Sepolia is on (EIP-7883 modexp pricing counts here).
evm_version = "osaka"
optimizer = true
optimizer_runs = 200
via_ir = true
fs_permissions = [{ access = "read", path = "./test/vectors" }, { access = "read-write", path = "./deploy-out.json" }]
auto_detect_remappings = false
[rpc_endpoints]
sepolia = "https://ethereum-sepolia-rpc.publicnode.com"

View file

@ -0,0 +1,46 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {Vm, VM_ADDRESS} from "../test/Vm.sol";
import {IgneumCertificateVerifier} from "../src/IgneumCertificateVerifier.sol";
/// Deploys the verifier on Sepolia from BRIDGE_DEPLOYER_KEY (environment, never printed), installs the voter table
/// from the vectors file named in BRIDGE_TABLE_JSON (a gen.mjs output: keys, weights, index, chain_id) and, when the
/// file carries a certificate (bitmap, signature), submits it, so the contract records one checkpoint final from the start.
///
/// BRIDGE_TABLE_JSON=test/vectors/chain.json forge script script/Deploy.s.sol:Deploy --rpc-url sepolia --broadcast --sig "run()"
contract Deploy {
Vm constant vm = Vm(VM_ADDRESS);
function run() external {
uint256 key = vm.envUint("BRIDGE_DEPLOYER_KEY");
string memory j = vm.readFile(vm.envOr("BRIDGE_TABLE_JSON", "test/vectors/chain.json"));
bytes[] memory keys = vm.parseJsonBytesArray(j, ".keys");
uint256[] memory w = vm.parseJsonUintArray(j, ".weights");
bytes memory packed;
uint64[] memory weights = new uint64[](w.length);
for (uint256 i = 0; i < keys.length; i++) {
packed = abi.encodePacked(packed, keys[i]);
weights[i] = uint64(w[i]);
}
uint64 index = uint64(vm.parseJsonUint(j, ".index"));
vm.startBroadcast(key);
IgneumCertificateVerifier v = new IgneumCertificateVerifier(vm.parseJsonString(j, ".chain_id"));
v.installTable(index, packed, weights);
bool hasCert = vm.keyExistsJson(j, ".bitmap");
if (hasCert) {
v.submitCertificate(index, vm.parseJsonBytes32(j, ".checkpoint"), vm.parseJsonBytes(j, ".bitmap"), vm.parseJsonBytes(j, ".signature"));
}
vm.stopBroadcast();
vm.writeFile(
"deploy-out.json",
string.concat(
"{\n \"IgneumCertificateVerifier\": \"", vm.toString(address(v)), "\",\n \"chain_id\": \"", vm.parseJsonString(j, ".chain_id"),
"\",\n \"table_index\": ", vm.toString(uint256(index)), ",\n \"voters\": ", vm.toString(keys.length), ",\n \"table_id\": \"",
vm.toString(v.tableId()), "\",\n \"final_checkpoint\": \"", hasCert ? vm.toString(vm.parseJsonBytes32(j, ".checkpoint")) : "none yet", "\"\n}\n"
)
);
}
}

View file

@ -0,0 +1,22 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {Vm, VM_ADDRESS} from "../test/Vm.sol";
/// Sends SEND_WEI of the chain's coin from the key in BRIDGE_DEPLOYER_KEY to SEND_TO (Sepolia test ETH between the
/// lanes' throwaway deployers). The key is read from the environment and never printed.
///
/// SEND_TO=0x.. SEND_WEI=20000000000000000 forge script script/Send.s.sol:Send --rpc-url sepolia --broadcast --sig "run()"
contract Send {
Vm constant vm = Vm(VM_ADDRESS);
function run() external {
uint256 key = vm.envUint("BRIDGE_DEPLOYER_KEY");
address to = vm.envAddress("SEND_TO");
uint256 wei_ = vm.envUint("SEND_WEI");
vm.startBroadcast(key);
(bool ok,) = to.call{value: wei_}("");
require(ok, "send failed");
vm.stopBroadcast();
}
}

View file

@ -0,0 +1,113 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
/// BLS12-381 through the EIP-2537 precompiles (Ethereum mainnet and Sepolia since Pectra): the hash-to-curve of
/// RFC 9380 (BLS12381G2_XMD:SHA-256_SSWU_RO_) with the caller's domain separation tag, public-key aggregation in G1
/// and the two-pairing check of a "minimal public key" signature (keys in G1, signatures in G2), the scheme of
/// Igneum's finality votes (consensus/core/src/finality.rs, blst "min_pk").
///
/// Encodings are the precompiles' own: a field element is 64 bytes (16 zero bytes then the 48-byte big-endian
/// value), a G1 point 128 bytes (x, y), a G2 point 256 bytes (x.c0, x.c1, y.c0, y.c1). Compressed chain forms
/// (48-byte keys, 96-byte signatures) are decompressed off chain by the submitter; the pairing precompile refuses
/// a point off the curve or outside the prime-order subgroup, so a wrong decompression fails the check.
library BLS12381 {
address internal constant G1ADD = address(0x0b);
address internal constant G2ADD = address(0x0d);
address internal constant PAIRING = address(0x0f);
address internal constant MAP_FP2_TO_G2 = address(0x11);
address internal constant MODEXP = address(0x05);
uint256 internal constant G1_LEN = 128;
uint256 internal constant G2_LEN = 256;
/// The field modulus p, big-endian, 48 bytes (the modexp precompile's modulus).
bytes internal constant P = hex"1a0111ea397fe69a4b1ba7b6434bacd764774b84f38512bf6730d2a0f6b0f6241eabfffeb153ffffb9feffffffffaaab";
/// The G1 generator with its y negated (p - y), in the 128-byte encoding, for the pairing check
/// e(pk, H(m)) * e(-G1, sig) == 1.
bytes internal constant NEG_G1 =
hex"0000000000000000000000000000000017f1d3a73197d7942695638c4fa9ac0fc3688c4f9774b905a14e3a3f171bac586c55e83ff97a1aeffb3af00adb22c6bb"
hex"00000000000000000000000000000000114d1d6855d545a8aa7d76c8cf2e21f267816aef1db507c96655b9d5caac42364e6f38ba0ecb751bad54dcd6b939c2ca";
error PrecompileFailed(address which);
error BadLength(string what);
// ---- hash to curve ----
/// expand_message_xmd with SHA-256 (RFC 9380 section 5.3.1) to `len` bytes; len at most 255 * 32.
function expandMessageXmd(bytes memory msg_, bytes memory dst, uint256 len) internal pure returns (bytes memory out) {
require(dst.length <= 255, "BLS: DST too long");
uint256 ell = (len + 31) / 32;
require(ell <= 255 && len > 0, "BLS: bad length");
bytes memory dstPrime = abi.encodePacked(dst, uint8(dst.length));
bytes32 b0 = sha256(abi.encodePacked(new bytes(64), msg_, uint16(len), uint8(0), dstPrime));
bytes32 bi = sha256(abi.encodePacked(b0, uint8(1), dstPrime));
out = new bytes(ell * 32);
assembly {
mstore(add(out, 32), bi)
}
for (uint256 i = 2; i <= ell; i++) {
bi = sha256(abi.encodePacked(b0 ^ bi, uint8(i), dstPrime));
assembly {
mstore(add(add(out, 32), mul(sub(i, 1), 32)), bi)
}
}
assembly {
mstore(out, len)
}
}
/// A 64-byte big-endian integer reduced mod p and returned in the precompiles' 64-byte field encoding.
function reduce64(bytes memory chunk, uint256 offset) internal view returns (bytes memory fe) {
require(chunk.length >= offset + 64, "BLS: chunk");
bytes memory base = new bytes(64);
for (uint256 i = 0; i < 64; i++) {
base[i] = chunk[offset + i];
}
// modexp(base^1 mod p): lengths 64, 1, 48
bytes memory input = abi.encodePacked(uint256(64), uint256(1), uint256(48), base, uint8(1), P);
(bool ok, bytes memory r) = MODEXP.staticcall(input);
if (!ok || r.length != 48) revert PrecompileFailed(MODEXP);
fe = abi.encodePacked(bytes16(0), r);
}
/// hash_to_curve for G2: two field elements of Fp2 from a 256-byte expansion, each mapped by the precompile
/// (which clears the cofactor), then added.
function hashToG2(bytes memory msg_, bytes memory dst) internal view returns (bytes memory point) {
bytes memory u = expandMessageXmd(msg_, dst, 256);
bytes memory q0 = mapFp2ToG2(abi.encodePacked(reduce64(u, 0), reduce64(u, 64)));
bytes memory q1 = mapFp2ToG2(abi.encodePacked(reduce64(u, 128), reduce64(u, 192)));
point = g2Add(q0, q1);
}
function mapFp2ToG2(bytes memory fp2) internal view returns (bytes memory point) {
if (fp2.length != 128) revert BadLength("fp2");
(bool ok, bytes memory r) = MAP_FP2_TO_G2.staticcall(fp2);
if (!ok || r.length != G2_LEN) revert PrecompileFailed(MAP_FP2_TO_G2);
point = r;
}
// ---- group operations ----
function g1Add(bytes memory a, bytes memory b) internal view returns (bytes memory c) {
if (a.length != G1_LEN || b.length != G1_LEN) revert BadLength("g1");
(bool ok, bytes memory r) = G1ADD.staticcall(abi.encodePacked(a, b));
if (!ok || r.length != G1_LEN) revert PrecompileFailed(G1ADD);
c = r;
}
function g2Add(bytes memory a, bytes memory b) internal view returns (bytes memory c) {
if (a.length != G2_LEN || b.length != G2_LEN) revert BadLength("g2");
(bool ok, bytes memory r) = G2ADD.staticcall(abi.encodePacked(a, b));
if (!ok || r.length != G2_LEN) revert PrecompileFailed(G2ADD);
c = r;
}
/// e(pk, hm) * e(-G1, sig) == 1, which holds exactly when sig = sk * hm for pk = sk * G1.
function verifyMinPk(bytes memory pk, bytes memory hm, bytes memory sig) internal view returns (bool) {
if (pk.length != G1_LEN || hm.length != G2_LEN || sig.length != G2_LEN) revert BadLength("pairing");
(bool ok, bytes memory r) = PAIRING.staticcall(abi.encodePacked(pk, hm, NEG_G1, sig));
if (!ok || r.length != 32) return false;
return abi.decode(r, (uint256)) == 1;
}
}

View file

@ -0,0 +1,155 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {BLS12381} from "./BLS12381.sol";
import {MerklePatricia} from "./MerklePatricia.sol";
interface IIgneumCertificateVerifier {
function chainId() external view returns (string memory);
function tableId() external view returns (bytes32);
function verifyCertificate(uint64 index, bytes32 checkpoint, bytes calldata bitmap, bytes calldata signature)
external
view
returns (bool ok, uint256 signedWeight, uint256 totalWeight);
function submitCertificate(uint64 index, bytes32 checkpoint, bytes calldata bitmap, bytes calldata signature) external;
function finalCheckpoint(uint64 index) external view returns (bytes32);
function isFinal(bytes32 checkpoint) external view returns (bool);
function verifyAccount(bytes32 stateRoot, address account, bytes[] calldata proof)
external
pure
returns (bool exists, uint256 nonce, uint256 balance, bytes32 storageRoot, bytes32 codeHash);
}
/// The Igneum light-client bridge primitive on Ethereum: verifies a Devnet 3 finality certificate (the aggregate
/// BLS signature of the canonical voter list over the vote message, under the 2/3-of-total-weight rule) against an
/// installed voter table, records the checkpoint hashes it proved final, and verifies an Ethereum-shape account
/// proof against a state root. What it proves and what it does not: docs/bridge/light-client-bridge.md.
///
/// Devnet 3, test tokens, no value.
contract IgneumCertificateVerifier is IIgneumCertificateVerifier {
using MerklePatricia for bytes32;
string public constant VOTE_PREFIX = "igneum-vote-v1/";
bytes public constant DST_VOTE = "IGNEUM_VOTE_V1_BLS12381G2_XMD:SHA-256_SSWU_RO_NUL_";
string private _chainId;
address public owner;
/// The canonical voter list at the installed checkpoint index: 128-byte G1 keys in the node's canonical order
/// (sorted by key hash, every key above dust and not stripped) and their weights (blue blocks in the window).
bytes[] private _keys;
uint64[] private _weights;
uint256 public totalWeight;
uint64 public tableIndex;
bytes32 public override tableId;
mapping(uint64 => bytes32) public override finalCheckpoint;
mapping(bytes32 => bool) public override isFinal;
event TableInstalled(uint64 indexed atIndex, uint256 voters, uint256 totalWeight, bytes32 tableId);
event CheckpointFinal(uint64 indexed index, bytes32 checkpoint, uint256 signedWeight, uint256 totalWeight, uint256 signers);
error NotOwner();
error NoTable();
error BadCertificate(string why);
constructor(string memory chainId_) {
_chainId = chainId_;
owner = msg.sender;
}
function chainId() external view override returns (string memory) {
return _chainId;
}
function voterCount() external view returns (uint256) {
return _keys.length;
}
function voter(uint256 i) external view returns (bytes memory key, uint64 weight) {
return (_keys[i], _weights[i]);
}
/// Installs the voter table read from a Devnet 3 node (igneum_getFinalityWeights at `atIndex`): `keys` is the
/// concatenation of 128-byte uncompressed G1 keys in canonical order, `weights` their weights. The table is a
/// trusted input of this first version (see the doc); only the installer may replace it.
function installTable(uint64 atIndex, bytes calldata keys, uint64[] calldata weights) external {
if (msg.sender != owner) revert NotOwner();
if (keys.length != weights.length * BLS12381.G1_LEN || weights.length == 0) revert BadCertificate("table shape");
delete _keys;
delete _weights;
uint256 total;
for (uint256 i = 0; i < weights.length; i++) {
_keys.push(keys[i * BLS12381.G1_LEN:(i + 1) * BLS12381.G1_LEN]);
_weights.push(weights[i]);
total += weights[i];
}
totalWeight = total;
tableIndex = atIndex;
tableId = keccak256(abi.encodePacked(atIndex, keys, abi.encodePacked(weights)));
emit TableInstalled(atIndex, weights.length, total, tableId);
}
/// The bytes every voter signs for (index, checkpoint): "igneum-vote-v1/" chain_id 0x00 index_le64 checkpoint.
function voteMessage(uint64 index, bytes32 checkpoint) public view returns (bytes memory) {
return abi.encodePacked(VOTE_PREFIX, _chainId, bytes1(0), le64(index), checkpoint);
}
function verifyCertificate(uint64 index, bytes32 checkpoint, bytes calldata bitmap, bytes calldata signature)
public
view
override
returns (bool ok, uint256 signedWeight, uint256 totalWeight_)
{
uint256 n = _keys.length;
if (n == 0) revert NoTable();
if (bitmap.length != (n + 7) / 8) revert BadCertificate("bitmap length");
if (signature.length != BLS12381.G2_LEN) revert BadCertificate("signature length");
bytes memory agg;
uint256 signers;
for (uint256 p = 0; p < n; p++) {
if (uint8(bitmap[p >> 3]) & uint8(1 << (p & 7)) == 0) continue;
signedWeight += _weights[p];
signers++;
agg = agg.length == 0 ? _keys[p] : BLS12381.g1Add(agg, _keys[p]);
}
totalWeight_ = totalWeight;
if (signers == 0) return (false, 0, totalWeight_);
// the rule decided 4 October 2026: signed weight at least two thirds of the whole window's weight
if (3 * signedWeight < 2 * totalWeight_) return (false, signedWeight, totalWeight_);
bytes memory hm = BLS12381.hashToG2(voteMessage(index, checkpoint), DST_VOTE);
ok = BLS12381.verifyMinPk(agg, hm, signature);
}
function submitCertificate(uint64 index, bytes32 checkpoint, bytes calldata bitmap, bytes calldata signature) external override {
(bool ok, uint256 signed, uint256 total) = verifyCertificate(index, checkpoint, bitmap, signature);
if (!ok) revert BadCertificate("certificate does not verify");
bytes32 known = finalCheckpoint[index];
if (known != bytes32(0) && known != checkpoint) revert BadCertificate("a different checkpoint is final at this index");
finalCheckpoint[index] = checkpoint;
isFinal[checkpoint] = true;
uint256 signers;
for (uint256 p = 0; p < _keys.length; p++) {
if (uint8(bitmap[p >> 3]) & uint8(1 << (p & 7)) != 0) signers++;
}
emit CheckpointFinal(index, checkpoint, signed, total, signers);
}
function verifyAccount(bytes32 stateRoot, address account, bytes[] calldata proof)
external
pure
override
returns (bool exists, uint256 nonce, uint256 balance, bytes32 storageRoot, bytes32 codeHash)
{
MerklePatricia.Account memory a = MerklePatricia.verifyAccount(stateRoot, account, proof);
return (a.exists, a.nonce, a.balance, a.storageRoot, a.codeHash);
}
function le64(uint64 v) internal pure returns (bytes8 out) {
uint64 r;
for (uint256 i = 0; i < 8; i++) {
r = (r << 8) | ((v >> (8 * i)) & 0xff);
}
out = bytes8(r);
}
}

View file

@ -0,0 +1,214 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
/// An Ethereum account proof (the eth_getProof shape) checked against a state root: the keccak-keyed Merkle
/// Patricia trie of reth's layout, which Igneum's executor uses for its stateRoot (igneum/exec/src/state.rs,
/// alloy_trie::root::state_root over keccak256(address) keys and RLP(nonce, balance, storageRoot, codeHash)
/// values). A proof is the list of RLP nodes from the root to the account's leaf, or to the branch or leaf that
/// shows the account absent.
library MerklePatricia {
struct Account {
bool exists;
uint256 nonce;
uint256 balance;
bytes32 storageRoot;
bytes32 codeHash;
}
error BadProof(string why);
/// Verifies `proof` for `account` under `stateRoot`; reverts when a node does not hash to its reference or
/// the path is malformed, returns exists=false when the trie shows no such account.
function verifyAccount(bytes32 stateRoot, address account, bytes[] memory proof) internal pure returns (Account memory out) {
bytes memory value = verifyPath(stateRoot, abi.encodePacked(keccak256(abi.encodePacked(account))), proof);
if (value.length == 0) return out;
(uint256 off, uint256 len, bool isList) = decode(value, 0);
if (!isList) revert BadProof("account value is not a list");
uint256 end = off + len;
uint256 p = off;
(uint256 o1, uint256 l1,) = decode(value, p);
out.nonce = toUint(value, o1, l1);
p = o1 + l1;
(uint256 o2, uint256 l2,) = decode(value, p);
out.balance = toUint(value, o2, l2);
p = o2 + l2;
(uint256 o3, uint256 l3,) = decode(value, p);
if (l3 != 32) revert BadProof("storage root length");
out.storageRoot = toBytes32(value, o3);
p = o3 + l3;
(uint256 o4, uint256 l4,) = decode(value, p);
if (l4 != 32) revert BadProof("code hash length");
out.codeHash = toBytes32(value, o4);
if (o4 + l4 != end) revert BadProof("account value has extra fields");
out.exists = true;
}
/// Walks the proof for `key` (32 bytes, hashed already) from `root`; returns the value found, or empty bytes
/// when the trie proves the key absent.
function verifyPath(bytes32 root, bytes memory key, bytes[] memory proof) internal pure returns (bytes memory value) {
bytes memory nibbles = toNibbles(key);
uint256 pos = 0;
bytes32 want = root;
bytes memory embedded;
for (uint256 i = 0; i < proof.length; i++) {
bytes memory node = proof[i];
if (embedded.length != 0) {
if (keccak256(node) != keccak256(embedded)) revert BadProof("embedded node mismatch");
embedded = "";
} else if (keccak256(node) != want) {
revert BadProof("node hash mismatch");
}
(uint256 off, uint256 len, bool isList) = decode(node, 0);
if (!isList) revert BadProof("node is not a list");
uint256 count = itemCount(node, off, len);
if (count == 17) {
if (pos == nibbles.length) {
// the branch's own value slot
(uint256 vo, uint256 vl,) = itemAt(node, off, 16);
return slice(node, vo, vl);
}
uint8 nib = uint8(nibbles[pos]);
(uint256 co, uint256 cl, bool clist) = itemAt(node, off, nib);
if (cl == 0 && !clist) return ""; // empty slot: the key is absent
pos++;
if (clist) {
embedded = slice(node, co - headerLen(node, co, cl, true), cl + headerLen(node, co, cl, true));
} else {
if (cl != 32) revert BadProof("child reference length");
want = toBytes32(node, co);
}
} else if (count == 2) {
(uint256 po, uint256 pl,) = itemAt(node, off, 0);
(bytes memory path, bool isLeaf) = decodePath(slice(node, po, pl));
if (!matches(nibbles, pos, path)) return ""; // diverging path: the key is absent
pos += path.length;
(uint256 vo, uint256 vl, bool vlist) = itemAt(node, off, 1);
if (isLeaf) {
if (pos != nibbles.length) revert BadProof("leaf before the key's end");
return slice(node, vo, vl);
}
if (vlist) {
embedded = slice(node, vo - headerLen(node, vo, vl, true), vl + headerLen(node, vo, vl, true));
} else {
if (vl != 32) revert BadProof("extension reference length");
want = toBytes32(node, vo);
}
} else {
revert BadProof("node arity");
}
}
revert BadProof("proof ends before the key");
}
// ---- paths ----
function toNibbles(bytes memory key) internal pure returns (bytes memory n) {
n = new bytes(key.length * 2);
for (uint256 i = 0; i < key.length; i++) {
n[2 * i] = bytes1(uint8(key[i]) >> 4);
n[2 * i + 1] = bytes1(uint8(key[i]) & 0x0f);
}
}
/// Hex-prefix decoding of a leaf or extension path.
function decodePath(bytes memory hp) internal pure returns (bytes memory path, bool isLeaf) {
if (hp.length == 0) revert BadProof("empty path");
uint8 flag = uint8(hp[0]) >> 4;
isLeaf = flag >= 2;
bool odd = flag % 2 == 1;
uint256 n = (hp.length - 1) * 2 + (odd ? 1 : 0);
path = new bytes(n);
uint256 w = 0;
if (odd) path[w++] = bytes1(uint8(hp[0]) & 0x0f);
for (uint256 i = 1; i < hp.length; i++) {
path[w++] = bytes1(uint8(hp[i]) >> 4);
path[w++] = bytes1(uint8(hp[i]) & 0x0f);
}
}
function matches(bytes memory nibbles, uint256 pos, bytes memory path) internal pure returns (bool) {
if (pos + path.length > nibbles.length) return false;
for (uint256 i = 0; i < path.length; i++) {
if (nibbles[pos + i] != path[i]) return false;
}
return true;
}
// ---- RLP ----
/// The item at `p`: the offset of its payload, the payload length and whether it is a list.
function decode(bytes memory b, uint256 p) internal pure returns (uint256 off, uint256 len, bool isList) {
if (p >= b.length) revert BadProof("rlp out of range");
uint8 first = uint8(b[p]);
if (first < 0x80) return (p, 1, false);
if (first < 0xb8) return (p + 1, first - 0x80, false);
if (first < 0xc0) {
uint256 n = first - 0xb7;
return (p + 1 + n, readLen(b, p + 1, n), false);
}
if (first < 0xf8) return (p + 1, first - 0xc0, true);
uint256 m = first - 0xf7;
return (p + 1 + m, readLen(b, p + 1, m), true);
}
function headerLen(bytes memory b, uint256 off, uint256 len, bool isList) private pure returns (uint256) {
// the header length of an item whose payload starts at off: single bytes under 0x80 have none
if (!isList && len == 1 && uint8(b[off]) < 0x80) return 0;
if (len < 56) return 1;
uint256 n = 0;
uint256 l = len;
while (l > 0) {
n++;
l >>= 8;
}
return 1 + n;
}
function readLen(bytes memory b, uint256 p, uint256 n) private pure returns (uint256 len) {
if (n == 0 || n > 32 || p + n > b.length) revert BadProof("rlp length");
for (uint256 i = 0; i < n; i++) {
len = (len << 8) | uint8(b[p + i]);
}
}
function itemCount(bytes memory b, uint256 off, uint256 len) private pure returns (uint256 n) {
uint256 p = off;
uint256 end = off + len;
while (p < end) {
(uint256 o, uint256 l,) = decode(b, p);
p = o + l;
n++;
}
if (p != end) revert BadProof("rlp list overrun");
}
function itemAt(bytes memory b, uint256 off, uint256 index) private pure returns (uint256 o, uint256 l, bool isList) {
uint256 p = off;
for (uint256 i = 0; ; i++) {
(o, l, isList) = decode(b, p);
if (i == index) return (o, l, isList);
p = o + l;
}
}
function toUint(bytes memory b, uint256 off, uint256 len) private pure returns (uint256 v) {
if (len > 32) revert BadProof("integer too long");
for (uint256 i = 0; i < len; i++) {
v = (v << 8) | uint8(b[off + i]);
}
}
function toBytes32(bytes memory b, uint256 off) private pure returns (bytes32 v) {
assembly {
v := mload(add(add(b, 32), off))
}
}
function slice(bytes memory b, uint256 off, uint256 len) private pure returns (bytes memory out) {
if (off + len > b.length) revert BadProof("slice out of range");
out = new bytes(len);
for (uint256 i = 0; i < len; i++) {
out[i] = b[off + i];
}
}
}

View file

@ -0,0 +1,144 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {Vm, VM_ADDRESS} from "./Vm.sol";
import {IgneumCertificateVerifier} from "../src/IgneumCertificateVerifier.sol";
import {BLS12381} from "../src/BLS12381.sol";
/// The verifier on Foundry's Prague EVM (the EIP-2537 precompiles): the synthetic vectors made by
/// test/vectors/gen.mjs (five keys, a certificate by four of them, a small account trie) and, when present, a real
/// certificate from the chain (test/vectors/chain.json, as /api/checkpoint serves it, decompressed by gen.mjs).
contract VerifierTest {
Vm constant vm = Vm(VM_ADDRESS);
string json;
IgneumCertificateVerifier v;
function setUp() public {
json = vm.readFile("test/vectors/synthetic.json");
v = new IgneumCertificateVerifier(vm.parseJsonString(json, ".chain_id"));
_install(v, json);
}
function _install(IgneumCertificateVerifier target, string memory j) internal {
bytes[] memory keys = vm.parseJsonBytesArray(j, ".keys");
uint256[] memory w = vm.parseJsonUintArray(j, ".weights");
bytes memory packed;
uint64[] memory weights = new uint64[](w.length);
for (uint256 i = 0; i < keys.length; i++) {
packed = abi.encodePacked(packed, keys[i]);
weights[i] = uint64(w[i]);
}
target.installTable(uint64(vm.parseJsonUint(j, ".index")), packed, weights);
}
function test_vote_message_matches_the_node() public view {
bytes memory want = vm.parseJsonBytes(json, ".vote_message");
bytes memory got = v.voteMessage(uint64(vm.parseJsonUint(json, ".index")), vm.parseJsonBytes32(json, ".checkpoint"));
require(keccak256(want) == keccak256(got), "vote message");
}
function test_expand_message_xmd_known_answer() public view {
// RFC 9380 appendix K.1 (expand_message_xmd with SHA-256, DST "QUUX-V01-CS02-with-expander-SHA256-128"): the
// empty message at 32 bytes is the appendix's own first answer; the "abc" at 128 bytes answer comes from noble
bytes memory dst = bytes(vm.parseJsonString(json, ".xmd_dst"));
bytes memory out = BLS12381.expandMessageXmd("", dst, 32);
require(keccak256(out) == keccak256(hex"68a985b87eb6b46952128911f2a4412bbc302a9d759667f87f7a21d803f07235"), "xmd 32 (RFC)");
require(keccak256(out) == keccak256(vm.parseJsonBytes(json, ".xmd_empty_32")), "xmd 32 (noble)");
bytes memory out2 = BLS12381.expandMessageXmd("abc", dst, 128);
require(keccak256(out2) == keccak256(vm.parseJsonBytes(json, ".xmd_abc_128")), "xmd 128 (noble)");
}
function test_certificate_verifies() public view {
(bool ok, uint256 signed, uint256 total) = v.verifyCertificate(
uint64(vm.parseJsonUint(json, ".index")), vm.parseJsonBytes32(json, ".checkpoint"), vm.parseJsonBytes(json, ".bitmap"), vm.parseJsonBytes(json, ".signature")
);
require(ok, "certificate");
require(signed == vm.parseJsonUint(json, ".signed_weight") && total == vm.parseJsonUint(json, ".total_weight"), "weights");
}
function test_certificate_under_two_thirds_is_refused() public view {
(bool ok, uint256 signed,) = v.verifyCertificate(
uint64(vm.parseJsonUint(json, ".index")), vm.parseJsonBytes32(json, ".checkpoint"), vm.parseJsonBytes(json, ".weak_bitmap"), vm.parseJsonBytes(json, ".weak_signature")
);
require(!ok && signed * 3 < vm.parseJsonUint(json, ".total_weight") * 2, "weak certificate accepted");
}
function test_wrong_checkpoint_or_index_fails() public view {
bytes32 cp = vm.parseJsonBytes32(json, ".checkpoint");
uint64 index = uint64(vm.parseJsonUint(json, ".index"));
bytes memory bm = vm.parseJsonBytes(json, ".bitmap");
bytes memory sig = vm.parseJsonBytes(json, ".signature");
(bool ok1,,) = v.verifyCertificate(index, cp ^ bytes32(uint256(1)), bm, sig);
(bool ok2,,) = v.verifyCertificate(index + 1, cp, bm, sig);
require(!ok1 && !ok2, "forged certificate accepted");
// the right signers' weight with a bitmap naming a different signer set does not match the signature
bytes memory other = vm.parseJsonBytes(json, ".weak_bitmap");
other[0] = bytes1(uint8(other[0]) | 0x1f);
(bool ok3,,) = v.verifyCertificate(index, cp, other, sig);
require(!ok3, "wrong signer set accepted");
}
function test_submit_records_the_checkpoint() public {
bytes32 cp = vm.parseJsonBytes32(json, ".checkpoint");
uint64 index = uint64(vm.parseJsonUint(json, ".index"));
v.submitCertificate(index, cp, vm.parseJsonBytes(json, ".bitmap"), vm.parseJsonBytes(json, ".signature"));
require(v.isFinal(cp) && v.finalCheckpoint(index) == cp, "not recorded");
vm.expectRevert(abi.encodeWithSelector(IgneumCertificateVerifier.BadCertificate.selector, "certificate does not verify"));
v.submitCertificate(index, cp, vm.parseJsonBytes(json, ".weak_bitmap"), vm.parseJsonBytes(json, ".weak_signature"));
}
function test_account_proof_present_and_absent() public view {
bytes32 root = vm.parseJsonBytes32(json, ".state_root");
(bool exists, uint256 nonce, uint256 balance, bytes32 sroot, bytes32 chash) =
v.verifyAccount(root, vm.parseJsonAddress(json, ".account"), vm.parseJsonBytesArray(json, ".account_proof"));
require(exists, "account absent");
require(nonce == vm.parseJsonUint(json, ".account_nonce") && balance == vm.parseJsonUint(json, ".account_balance"), "account fields");
require(sroot == vm.parseJsonBytes32(json, ".account_storage_root") && chash == vm.parseJsonBytes32(json, ".account_code_hash"), "account roots");
(bool exists2,,,,) = v.verifyAccount(root, vm.parseJsonAddress(json, ".absent_account"), vm.parseJsonBytesArray(json, ".absent_proof"));
require(!exists2, "absent account present");
}
function test_account_proof_against_a_wrong_root_reverts() public {
bytes32 root = vm.parseJsonBytes32(json, ".state_root") ^ bytes32(uint256(1));
bytes[] memory proof = vm.parseJsonBytesArray(json, ".account_proof");
address a = vm.parseJsonAddress(json, ".account");
vm.expectRevert(abi.encodeWithSelector(bytes4(keccak256("BadProof(string)")), "node hash mismatch"));
v.verifyAccount(root, a, proof);
}
/// A certificate the chain actually carried (test/vectors/chain.json; skipped when the file is absent).
function test_chain_certificate_verifies() public {
string memory j;
try vm.readFile("test/vectors/chain.json") returns (string memory s) {
j = s;
} catch {
return;
}
IgneumCertificateVerifier c = new IgneumCertificateVerifier(vm.parseJsonString(j, ".chain_id"));
_install(c, j);
(bool ok, uint256 signed, uint256 total) = c.verifyCertificate(
uint64(vm.parseJsonUint(j, ".index")), vm.parseJsonBytes32(j, ".checkpoint"), vm.parseJsonBytes(j, ".bitmap"), vm.parseJsonBytes(j, ".signature")
);
require(signed == vm.parseJsonUint(j, ".signed_weight") && total == vm.parseJsonUint(j, ".total_weight"), "chain weights");
require(ok, "the chain's certificate does not verify");
}
/// One Devnet 3 account under a real Devnet 3 state root (test/vectors/dn3-account.json from a node's eth_getProof;
/// skipped when the file is absent). The root is the chain's own; the link from a certified checkpoint to that root
/// is the gap the doc names.
function test_devnet3_account_balance_proven() public {
string memory j;
try vm.readFile("test/vectors/dn3-account.json") returns (string memory s) {
j = s;
} catch {
return;
}
(bool exists, uint256 nonce, uint256 balance, bytes32 sroot, bytes32 chash) = v.verifyAccount(
vm.parseJsonBytes32(j, ".state_root"), vm.parseJsonAddress(j, ".account"), vm.parseJsonBytesArray(j, ".account_proof")
);
require(exists, "the Devnet 3 account is absent under its root");
require(nonce == vm.parseJsonUint(j, ".account_nonce") && balance == vm.parseJsonUint(j, ".account_balance"), "Devnet 3 account fields");
require(sroot == vm.parseJsonBytes32(j, ".account_storage_root") && chash == vm.parseJsonBytes32(j, ".account_code_hash"), "Devnet 3 account roots");
}
}

View file

@ -0,0 +1,34 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
/// The Foundry cheatcodes this project uses, declared here so the tree needs no remote dependency.
interface Vm {
function startBroadcast(uint256 privateKey) external;
function stopBroadcast() external;
function envUint(string calldata name) external view returns (uint256);
function envAddress(string calldata name) external view returns (address);
function envOr(string calldata name, string calldata defaultValue) external view returns (string memory);
function toString(address value) external pure returns (string memory);
function toString(uint256 value) external pure returns (string memory);
function toString(bytes32 value) external pure returns (string memory);
function toString(bytes calldata value) external pure returns (string memory);
function readFile(string calldata path) external view returns (string memory);
function writeFile(string calldata path, string calldata data) external;
function parseJsonBytes(string calldata json, string calldata key) external pure returns (bytes memory);
function parseJsonBytes32(string calldata json, string calldata key) external pure returns (bytes32);
function parseJsonUint(string calldata json, string calldata key) external pure returns (uint256);
function parseJsonString(string calldata json, string calldata key) external pure returns (string memory);
function parseJsonBytesArray(string calldata json, string calldata key) external pure returns (bytes[] memory);
function parseJsonUintArray(string calldata json, string calldata key) external pure returns (uint256[] memory);
function parseJsonAddress(string calldata json, string calldata key) external pure returns (address);
function keyExistsJson(string calldata json, string calldata key) external view returns (bool);
function deal(address who, uint256 newBalance) external;
function prank(address msgSender) external;
function startPrank(address msgSender) external;
function stopPrank() external;
function warp(uint256 newTimestamp) external;
function expectRevert(bytes calldata revertData) external;
function addr(uint256 privateKey) external pure returns (address);
}
address constant VM_ADDRESS = address(uint160(uint256(keccak256("hevm cheat code"))));

View file

@ -0,0 +1,9 @@
node_modules
synthetic.json
chain.json
checkpoint-live.json
checkpoint-dn3.json
weights-dn3.json
checkpoints-dn3.json
dn3-table.json
dn3-account.json

View file

@ -0,0 +1,116 @@
// Test vectors for the Igneum certificate verifier. Two sources:
// node gen.mjs synthetic > synthetic.json five vote keys made here (noble BLS12-381), a certificate signed by four
// of them over the Devnet 3 vote message, a small account trie with proofs
// node gen.mjs table <weights.json> > dn3-table.json the voter table alone from a node's igneum_getFinalityWeights answer
// (voters in the canonical order: sorted by key hash; weight = blocks)
// node gen.mjs account <getProof.json> <stateRoot> <blockNumber> > dn3-account.json an eth_getProof answer from a Devnet 3
// node (the reference-apps lane's reader) as the suite's real-root account vector
// node gen.mjs chain <checkpoint.json> > dn3.json a real certificate as igneum.network/api/checkpoint?source=dn3 serves it:
// the voter table and the aggregate signature decompressed to the
// precompiles' encodings (the verifier checks the same bytes the node signed)
// Encodings: field element 64 bytes (16 zero bytes then 48), G1 128 bytes, G2 256 bytes (x.c0, x.c1, y.c0, y.c1).
// The DST and the vote message follow consensus/core/src/finality.rs and site/verify/core.js.
import { bls12_381 } from '@noble/curves/bls12-381';
import { expand_message_xmd } from '@noble/curves/abstract/hash-to-curve';
import { sha256 } from '@noble/hashes/sha256';
import { readFileSync } from 'node:fs';
import { createRequire } from 'node:module';
const require = createRequire(import.meta.url);
const DST = 'IGNEUM_VOTE_V1_BLS12381G2_XMD:SHA-256_SSWU_RO_NUL_';
const CHAIN_ID = 'igneum-devnet-3';
const te = new TextEncoder();
const hex = b => '0x' + Array.from(b, x => x.toString(16).padStart(2, '0')).join('');
const unhex = h => Uint8Array.from(Buffer.from(h.replace(/^0x/, ''), 'hex'));
const be = (n, len) => { const out = new Uint8Array(len); let v = BigInt(n); for (let i = len - 1; i >= 0; i--) { out[i] = Number(v & 0xffn); v >>= 8n; } return out; };
const fe = n => { const out = new Uint8Array(64); out.set(be(n, 48), 16); return out; };
const concat = parts => { const n = parts.reduce((a, p) => a + p.length, 0); const out = new Uint8Array(n); let o = 0; for (const p of parts) { out.set(p, o); o += p.length; } return out; };
const u64le = n => { const out = new Uint8Array(8); let v = BigInt(n); for (let i = 0; i < 8; i++) { out[i] = Number(v & 0xffn); v >>= 8n; } return out; };
const G1 = bls12_381.G1.ProjectivePoint, G2 = bls12_381.G2.ProjectivePoint;
function g1Enc(p) { const a = p.toAffine(); return concat([fe(a.x), fe(a.y)]); }
function g2Enc(p) { const a = p.toAffine(); return concat([fe(a.x.c0), fe(a.x.c1), fe(a.y.c0), fe(a.y.c1)]); }
function voteMessage(index, checkpointHex) { return concat([te.encode('igneum-vote-v1/' + CHAIN_ID), new Uint8Array([0]), u64le(index), unhex(checkpointHex)]); }
function bitmapOf(positions, n) { const bm = new Uint8Array(Math.ceil(n / 8)); for (const p of positions) bm[p >> 3] |= 1 << (p & 7); return bm; }
async function synthetic() {
const sks = [1, 2, 3, 4, 5].map(i => { const s = new Uint8Array(32); s[31] = i; s[0] = 0x11 * i; return bls12_381.utils.randomPrivateKey ? bls12_381.G1.normPrivateKeyToScalar(s) : s; });
const keys = sks.map(sk => G1.BASE.multiply(sk));
const weights = [100, 250, 400, 300, 150];
const index = 1234, checkpoint = '0x' + 'ab'.repeat(32);
const msg = voteMessage(index, checkpoint);
const hm = bls12_381.G2.hashToCurve(msg, { DST });
const signers = [0, 1, 2, 3]; // 1,050 of 1,200: above two thirds
const weakSigners = [0, 2, 4]; // 650 of 1,200: under two thirds
const sign = who => who.map(i => hm.multiply(sks[i])).reduce((a, b) => a.add(b));
const sig = sign(signers), weak = sign(weakSigners);
// the account trie: three accounts, proofs for one present and one absent
const { Trie } = require('@ethereumjs/trie');
const { RLP } = require('@ethereumjs/rlp');
const { keccak256 } = require('ethereum-cryptography/keccak');
const trie = new Trie({ useKeyHashing: true });
const accounts = [
{ address: '0x07dd4dbca5c1a66755af28bacca1d901a2d209aa', nonce: 7n, balance: 999174011168718479514n, storageRoot: '0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421', codeHash: '0xc5d2460186f7233c927e7db2dcc703c0e500b653ca82273b7bfad8045d85a470' },
{ address: '0x9a6fa842c4e58a87aef1f3ad15233d99283002b7', nonce: 1n, balance: 0n, storageRoot: '0x' + '11'.repeat(32), codeHash: '0x' + '22'.repeat(32) },
{ address: '0x53fe98022c2ac26d5d721457fb1c374b4d56144b', nonce: 3n, balance: 2580n * 10n ** 18n, storageRoot: '0x56e81f171bcc55a6ff8345e692c0f86e5b48e01b996cadc001622fb5e363b421', codeHash: '0xc5d2460186f7233c927e7db2dcc703c0e500b653ca82273b7bfad8045d85a470' },
];
for (const a of accounts) {
const v = RLP.encode([a.nonce === 0n ? new Uint8Array() : be(a.nonce, Math.ceil(a.nonce.toString(2).length / 8)), a.balance === 0n ? new Uint8Array() : be(a.balance, Math.ceil(a.balance.toString(2).length / 8)), unhex(a.storageRoot), unhex(a.codeHash)]);
await trie.put(unhex(a.address), v);
}
const root = hex(trie.root());
const proofFor = async addr => (await trie.createProof(unhex(addr))).map(hex);
const absent = '0x000000000000000000000000000000000000dead';
return {
chain_id: CHAIN_ID, dst: DST, index, checkpoint,
keys: keys.map(k => hex(g1Enc(k))), weights, total_weight: weights.reduce((a, b) => a + b, 0),
bitmap: hex(bitmapOf(signers, keys.length)), signature: hex(g2Enc(sig)), signed_weight: signers.reduce((a, i) => a + weights[i], 0),
weak_bitmap: hex(bitmapOf(weakSigners, keys.length)), weak_signature: hex(g2Enc(weak)),
vote_message: hex(msg),
// RFC 9380 expand_message_xmd(SHA-256) answers, computed by noble, for the Solidity port's own check
xmd_dst: 'QUUX-V01-CS02-with-expander-SHA256-128',
xmd_abc_128: hex(expand_message_xmd(te.encode('abc'), te.encode('QUUX-V01-CS02-with-expander-SHA256-128'), 128, sha256)),
xmd_empty_32: hex(expand_message_xmd(new Uint8Array(), te.encode('QUUX-V01-CS02-with-expander-SHA256-128'), 32, sha256)),
state_root: root,
account: accounts[0].address, account_nonce: accounts[0].nonce.toString(), account_balance: accounts[0].balance.toString(),
account_storage_root: accounts[0].storageRoot, account_code_hash: accounts[0].codeHash,
account_proof: await proofFor(accounts[0].address),
absent_account: absent, absent_proof: await proofFor(absent),
};
}
function chain(file) {
const d = JSON.parse(readFileSync(file, 'utf8'));
const voters = d.voters.map(v => ({ key: hex(g1Enc(G1.fromHex(v.pubkey_hex.replace(/^0x/, '')))), weight: Math.round(Number(v.weight)) }));
const sig = G2.fromHex(d.certificate.aggregate_signature_hex.replace(/^0x/, ''));
const positions = []; const bm = unhex(d.certificate.bitmap_hex);
for (let p = 0; p < voters.length; p++) if (bm[p >> 3] & (1 << (p & 7))) positions.push(p);
return {
source: d.source, chain_id: d.chain_id, dst: DST, index: d.index, checkpoint: '0x' + d.hash,
keys: voters.map(v => v.key), weights: voters.map(v => v.weight), total_weight: voters.reduce((a, v) => a + v.weight, 0),
bitmap: '0x' + d.certificate.bitmap_hex, signature: hex(g2Enc(sig)), signed_weight: positions.reduce((a, p) => a + voters[p].weight, 0),
signers: positions.length, voters_at_index: d.voters_at_index, stored_at: d.stored_at,
};
}
function table(file) {
const d = JSON.parse(readFileSync(file, 'utf8'));
const r = d.result || d;
const voters = r.keys.filter(k => k.voter === true || k.voter === 'True').map(k => ({ keyHash: String(k.keyHash).replace(/^0x/, ''), key: hex(g1Enc(G1.fromHex(String(k.pubkey).replace(/^0x/, '')))), weight: Number(BigInt(k.blocks)) }));
voters.sort((a, b) => (a.keyHash < b.keyHash ? -1 : a.keyHash > b.keyHash ? 1 : 0));
const total = voters.reduce((a, v) => a + v.weight, 0);
return { chain_id: process.env.IGNEUM_CHAIN_ID || CHAIN_ID, dst: DST, index: Number(BigInt(r.checkpointIndex)), checkpoint_at_index: '0x' + String(r.checkpointHash).replace(/^0x/, ''), keys: voters.map(v => v.key), weights: voters.map(v => v.weight), total_weight: total, total_weight_node: Number(BigInt(r.totalWeight)), voters: voters.length };
}
function account(file, stateRoot, blockNumber) {
const d = JSON.parse(readFileSync(file, 'utf8'));
const r = d.result || d;
return { chain_id: CHAIN_ID, state_root: stateRoot, block_number: Number(blockNumber), account: r.address, account_nonce: BigInt(r.nonce).toString(), account_balance: BigInt(r.balance).toString(), account_storage_root: r.storageHash, account_code_hash: r.codeHash, account_proof: r.accountProof };
}
const mode = process.argv[2];
if (mode === 'synthetic') synthetic().then(v => console.log(JSON.stringify(v, null, 1)));
else if (mode === 'chain') console.log(JSON.stringify(chain(process.argv[3]), null, 1));
else if (mode === 'table') console.log(JSON.stringify(table(process.argv[3]), null, 1));
else if (mode === 'account') console.log(JSON.stringify(account(process.argv[3], process.argv[4], process.argv[5]), null, 1));
else { console.error('usage: gen.mjs synthetic | table <weights.json> | account <getProof.json> <stateRoot> <blockNumber> | chain <checkpoint.json>'); process.exit(2); }

View file

@ -0,0 +1,19 @@
[profile.default]
src = "src"
test = "test"
script = "script"
out = "out"
libs = []
solc_version = "0.8.28"
# Devnet 3's executor runs the EVM through revm; Paris keeps the bytecode off PUSH0 and transient storage so it runs
# on any post-Merge configuration of it (8 October 2026).
evm_version = "paris"
optimizer = true
optimizer_runs = 200
fs_permissions = [{ access = "read-write", path = "./deploy-out.json" }]
# no remote dependencies: the tests and the script declare the cheatcode interface they use (test/Vm.sol); solc 0.8.28 is fetched once by forge
auto_detect_remappings = false
[rpc_endpoints]
# the Devnet 3 EVM reaches the build box through a tunnel on this port (never the Mac); see docs/contracts/devnet-3.json
devnet3 = "http://127.0.0.1:36790"

View file

@ -0,0 +1,58 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {Vm, VM_ADDRESS} from "../test/Vm.sol";
import {WIGN} from "../src/WIGN.sol";
import {TestToken} from "../src/TestToken.sol";
import {IgneumFactory} from "../src/IgneumPair.sol";
import {IgneumRouter} from "../src/IgneumRouter.sol";
/// Deploys the AMM on Devnet 3 from the key in DEX_DEPLOYER_KEY (read from the environment, never printed), seeds
/// three pools from the faucet and 200 IGN, and writes the addresses to deploy-out.json for docs/contracts.
///
/// forge script script/Deploy.s.sol:Deploy --rpc-url devnet3 --broadcast --sig "run()"
///
/// Devnet 3, test tokens, no value.
contract Deploy {
Vm constant vm = Vm(VM_ADDRESS);
function run() external {
uint256 key = vm.envUint("DEX_DEPLOYER_KEY");
address deployer = vm.addr(key);
vm.startBroadcast(key);
WIGN wign = new WIGN();
TestToken tta = new TestToken("Test Token A", "TTA");
TestToken ttb = new TestToken("Test Token B", "TTB");
IgneumFactory factory = new IgneumFactory();
IgneumRouter router = new IgneumRouter(address(factory), address(wign));
tta.drip();
ttb.drip();
tta.approve(address(router), type(uint256).max);
ttb.approve(address(router), type(uint256).max);
uint256 deadline = block.timestamp + 1 hours;
router.addLiquidityIGN{value: 100 ether}(address(tta), 500 ether, 0, 0, deployer, deadline);
router.addLiquidityIGN{value: 100 ether}(address(ttb), 500 ether, 0, 0, deployer, deadline);
router.addLiquidity(address(tta), address(ttb), 500 ether, 500 ether, 0, 0, deployer, deadline);
vm.stopBroadcast();
string memory json = "{\n";
json = _line(json, "WIGN", address(wign), ",");
json = _line(json, "TTA", address(tta), ",");
json = _line(json, "TTB", address(ttb), ",");
json = _line(json, "IgneumFactory", address(factory), ",");
json = _line(json, "IgneumRouter", address(router), ",");
json = _line(json, "pair_TTA_WIGN", factory.getPair(address(tta), address(wign)), ",");
json = _line(json, "pair_TTB_WIGN", factory.getPair(address(ttb), address(wign)), ",");
json = _line(json, "pair_TTA_TTB", factory.getPair(address(tta), address(ttb)), ",");
json = _line(json, "deployer", deployer, "");
json = string.concat(json, "}\n");
vm.writeFile("deploy-out.json", json);
}
function _line(string memory json, string memory key, address value, string memory comma) private pure returns (string memory) {
return string.concat(json, " \"", key, "\": \"", vm.toString(value), "\"", comma, "\n");
}
}

View file

@ -0,0 +1,37 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {Vm, VM_ADDRESS} from "../test/Vm.sol";
import {TestToken} from "../src/TestToken.sol";
import {IgneumRouter} from "../src/IgneumRouter.sol";
/// Seeds the three pools of an already deployed AMM (the addresses from the environment): a drip of each test
/// token, approvals, 100 IGN beside 500 TTA, 100 IGN beside 500 TTB, 500 TTA beside 500 TTB. Run with
/// --skip-simulation so every gas limit comes from the node's own eth_estimateGas: Devnet 3 charges the proving
/// dimension inside the execution gas, so Foundry's local estimate runs out (8 October 2026, drip() at 133,603 gas:
/// out of gas; the node's estimate 685,513).
///
/// DEX_TTA=0x.. DEX_TTB=0x.. DEX_ROUTER=0x.. forge script script/Seed.s.sol:Seed --rpc-url devnet3 --broadcast --slow --skip-simulation --sig "run()"
///
/// Devnet 3, test tokens, no value.
contract Seed {
Vm constant vm = Vm(VM_ADDRESS);
function run() external {
uint256 key = vm.envUint("DEX_DEPLOYER_KEY");
address deployer = vm.addr(key);
TestToken tta = TestToken(vm.envAddress("DEX_TTA"));
TestToken ttb = TestToken(vm.envAddress("DEX_TTB"));
IgneumRouter router = IgneumRouter(payable(vm.envAddress("DEX_ROUTER")));
vm.startBroadcast(key);
if (tta.dripWait(deployer) == 0 && tta.balanceOf(deployer) < 1_000 ether) tta.drip();
if (ttb.dripWait(deployer) == 0 && ttb.balanceOf(deployer) < 1_000 ether) ttb.drip();
tta.approve(address(router), type(uint256).max);
ttb.approve(address(router), type(uint256).max);
uint256 deadline = block.timestamp + 1 hours;
router.addLiquidityIGN{value: 100 ether}(address(tta), 500 ether, 0, 0, deployer, deadline);
router.addLiquidityIGN{value: 100 ether}(address(ttb), 500 ether, 0, 0, deployer, deadline);
router.addLiquidity(address(tta), address(ttb), 500 ether, 500 ether, 0, 0, deployer, deadline);
vm.stopBroadcast();
}
}

View file

@ -0,0 +1,59 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
/// A plain ERC-20 with EIP-2612-free approvals, shared by the pair's liquidity token, the wrapped coin and the test
/// tokens. Devnet 3, test tokens, no value.
contract ERC20 {
string public name;
string public symbol;
uint8 public constant decimals = 18;
uint256 public totalSupply;
mapping(address => uint256) public balanceOf;
mapping(address => mapping(address => uint256)) public allowance;
event Transfer(address indexed from, address indexed to, uint256 value);
event Approval(address indexed owner, address indexed spender, uint256 value);
constructor(string memory name_, string memory symbol_) {
name = name_;
symbol = symbol_;
}
function _mint(address to, uint256 value) internal {
totalSupply += value;
balanceOf[to] += value;
emit Transfer(address(0), to, value);
}
function _burn(address from, uint256 value) internal {
balanceOf[from] -= value;
totalSupply -= value;
emit Transfer(from, address(0), value);
}
function _transfer(address from, address to, uint256 value) internal {
balanceOf[from] -= value;
balanceOf[to] += value;
emit Transfer(from, to, value);
}
function approve(address spender, uint256 value) external returns (bool) {
allowance[msg.sender][spender] = value;
emit Approval(msg.sender, spender, value);
return true;
}
function transfer(address to, uint256 value) external returns (bool) {
_transfer(msg.sender, to, value);
return true;
}
function transferFrom(address from, address to, uint256 value) external returns (bool) {
uint256 allowed = allowance[from][msg.sender];
if (allowed != type(uint256).max) {
allowance[from][msg.sender] = allowed - value;
}
_transfer(from, to, value);
return true;
}
}

View file

@ -0,0 +1,174 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {ERC20} from "./ERC20.sol";
interface IERC20Minimal {
function balanceOf(address) external view returns (uint256);
function transfer(address, uint256) external returns (bool);
}
/// A constant-product pool in the Uniswap v2 shape: reserves of two tokens, a 0.3 percent fee kept in the pool,
/// liquidity tokens for the depositors, MINIMUM_LIQUIDITY locked at the first mint. No protocol fee, no price
/// accumulators, no flash callback. Devnet 3, test tokens, no value.
contract IgneumPair is ERC20("Igneum LP", "IGN-LP") {
uint256 public constant MINIMUM_LIQUIDITY = 10 ** 3;
address public immutable factory;
address public token0;
address public token1;
uint112 private reserve0;
uint112 private reserve1;
uint32 private blockTimestampLast;
uint256 private unlocked = 1;
event Mint(address indexed sender, uint256 amount0, uint256 amount1);
event Burn(address indexed sender, uint256 amount0, uint256 amount1, address indexed to);
event Swap(address indexed sender, uint256 amount0In, uint256 amount1In, uint256 amount0Out, uint256 amount1Out, address indexed to);
event Sync(uint112 reserve0, uint112 reserve1);
modifier lock() {
require(unlocked == 1, "Pair: locked");
unlocked = 0;
_;
unlocked = 1;
}
constructor() {
factory = msg.sender;
}
function initialize(address token0_, address token1_) external {
require(msg.sender == factory, "Pair: forbidden");
token0 = token0_;
token1 = token1_;
}
function getReserves() public view returns (uint112, uint112, uint32) {
return (reserve0, reserve1, blockTimestampLast);
}
function _safeTransfer(address token, address to, uint256 value) private {
(bool ok, bytes memory data) = token.call(abi.encodeWithSelector(IERC20Minimal.transfer.selector, to, value));
require(ok && (data.length == 0 || abi.decode(data, (bool))), "Pair: transfer failed");
}
function _update(uint256 balance0, uint256 balance1) private {
require(balance0 <= type(uint112).max && balance1 <= type(uint112).max, "Pair: overflow");
reserve0 = uint112(balance0);
reserve1 = uint112(balance1);
blockTimestampLast = uint32(block.timestamp);
emit Sync(reserve0, reserve1);
}
function _sqrt(uint256 y) private pure returns (uint256 z) {
if (y > 3) {
z = y;
uint256 x = y / 2 + 1;
while (x < z) {
z = x;
x = (y / x + x) / 2;
}
} else if (y != 0) {
z = 1;
}
}
function _min(uint256 x, uint256 y) private pure returns (uint256) {
return x < y ? x : y;
}
/// Mints liquidity for the tokens sent to the pair since the last sync. Called by the router.
function mint(address to) external lock returns (uint256 liquidity) {
(uint112 r0, uint112 r1,) = getReserves();
uint256 balance0 = IERC20Minimal(token0).balanceOf(address(this));
uint256 balance1 = IERC20Minimal(token1).balanceOf(address(this));
uint256 amount0 = balance0 - r0;
uint256 amount1 = balance1 - r1;
if (totalSupply == 0) {
liquidity = _sqrt(amount0 * amount1) - MINIMUM_LIQUIDITY;
_mint(address(0xdead), MINIMUM_LIQUIDITY);
} else {
liquidity = _min(amount0 * totalSupply / r0, amount1 * totalSupply / r1);
}
require(liquidity > 0, "Pair: insufficient liquidity minted");
_mint(to, liquidity);
_update(balance0, balance1);
emit Mint(msg.sender, amount0, amount1);
}
/// Burns the liquidity tokens sent to the pair and pays both tokens out pro rata. Called by the router.
function burn(address to) external lock returns (uint256 amount0, uint256 amount1) {
uint256 balance0 = IERC20Minimal(token0).balanceOf(address(this));
uint256 balance1 = IERC20Minimal(token1).balanceOf(address(this));
uint256 liquidity = balanceOf[address(this)];
amount0 = liquidity * balance0 / totalSupply;
amount1 = liquidity * balance1 / totalSupply;
require(amount0 > 0 && amount1 > 0, "Pair: insufficient liquidity burned");
_burn(address(this), liquidity);
_safeTransfer(token0, to, amount0);
_safeTransfer(token1, to, amount1);
_update(IERC20Minimal(token0).balanceOf(address(this)), IERC20Minimal(token1).balanceOf(address(this)));
emit Burn(msg.sender, amount0, amount1, to);
}
/// Pays out up to the amounts asked and checks the fee-adjusted product did not fall. The input must already
/// sit in the pair (the router sends it first).
function swap(uint256 amount0Out, uint256 amount1Out, address to) external lock {
require(amount0Out > 0 || amount1Out > 0, "Pair: insufficient output amount");
(uint112 r0, uint112 r1,) = getReserves();
require(amount0Out < r0 && amount1Out < r1, "Pair: insufficient liquidity");
require(to != token0 && to != token1, "Pair: invalid to");
if (amount0Out > 0) _safeTransfer(token0, to, amount0Out);
if (amount1Out > 0) _safeTransfer(token1, to, amount1Out);
uint256 balance0 = IERC20Minimal(token0).balanceOf(address(this));
uint256 balance1 = IERC20Minimal(token1).balanceOf(address(this));
uint256 amount0In = balance0 > r0 - amount0Out ? balance0 - (r0 - amount0Out) : 0;
uint256 amount1In = balance1 > r1 - amount1Out ? balance1 - (r1 - amount1Out) : 0;
require(amount0In > 0 || amount1In > 0, "Pair: insufficient input amount");
uint256 adjusted0 = balance0 * 1000 - amount0In * 3;
uint256 adjusted1 = balance1 * 1000 - amount1In * 3;
require(adjusted0 * adjusted1 >= uint256(r0) * uint256(r1) * 1000 ** 2, "Pair: K");
_update(balance0, balance1);
emit Swap(msg.sender, amount0In, amount1In, amount0Out, amount1Out, to);
}
/// Sends any balance above the reserves to `to`.
function skim(address to) external lock {
_safeTransfer(token0, to, IERC20Minimal(token0).balanceOf(address(this)) - reserve0);
_safeTransfer(token1, to, IERC20Minimal(token1).balanceOf(address(this)) - reserve1);
}
/// Sets the reserves to the balances.
function sync() external lock {
_update(IERC20Minimal(token0).balanceOf(address(this)), IERC20Minimal(token1).balanceOf(address(this)));
}
}
/// Creates one pair per unordered token pair and remembers it.
contract IgneumFactory {
mapping(address => mapping(address => address)) public getPair;
address[] public allPairs;
event PairCreated(address indexed token0, address indexed token1, address pair, uint256 count);
function allPairsLength() external view returns (uint256) {
return allPairs.length;
}
function createPair(address tokenA, address tokenB) external returns (address pair) {
require(tokenA != tokenB, "Factory: identical addresses");
(address token0, address token1) = tokenA < tokenB ? (tokenA, tokenB) : (tokenB, tokenA);
require(token0 != address(0), "Factory: zero address");
require(getPair[token0][token1] == address(0), "Factory: pair exists");
bytes32 salt = keccak256(abi.encodePacked(token0, token1));
pair = address(new IgneumPair{salt: salt}());
IgneumPair(pair).initialize(token0, token1);
getPair[token0][token1] = pair;
getPair[token1][token0] = pair;
allPairs.push(pair);
emit PairCreated(token0, token1, pair, allPairs.length);
}
}

View file

@ -0,0 +1,253 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {IgneumFactory, IgneumPair} from "./IgneumPair.sol";
interface IERC20Router {
function balanceOf(address) external view returns (uint256);
function transfer(address, uint256) external returns (bool);
function transferFrom(address, address, uint256) external returns (bool);
}
interface IWIGN is IERC20Router {
function deposit() external payable;
function withdraw(uint256) external;
}
/// The router in the Uniswap v2 shape: adds and removes liquidity, swaps along a path of pairs, quotes. Pairs are
/// looked up on the factory, never derived from an init-code hash. Devnet 3, test tokens, no value.
contract IgneumRouter {
IgneumFactory public immutable factory;
address public immutable WIGN;
modifier ensure(uint256 deadline) {
require(deadline >= block.timestamp, "Router: expired");
_;
}
constructor(address factory_, address wign_) {
factory = IgneumFactory(factory_);
WIGN = wign_;
}
receive() external payable {
require(msg.sender == WIGN, "Router: only WIGN");
}
// ---- pure maths ----
function sortTokens(address tokenA, address tokenB) public pure returns (address token0, address token1) {
require(tokenA != tokenB, "Router: identical addresses");
(token0, token1) = tokenA < tokenB ? (tokenA, tokenB) : (tokenB, tokenA);
require(token0 != address(0), "Router: zero address");
}
function quote(uint256 amountA, uint256 reserveA, uint256 reserveB) public pure returns (uint256 amountB) {
require(amountA > 0, "Router: insufficient amount");
require(reserveA > 0 && reserveB > 0, "Router: insufficient liquidity");
amountB = amountA * reserveB / reserveA;
}
function getAmountOut(uint256 amountIn, uint256 reserveIn, uint256 reserveOut) public pure returns (uint256 amountOut) {
require(amountIn > 0, "Router: insufficient input amount");
require(reserveIn > 0 && reserveOut > 0, "Router: insufficient liquidity");
uint256 amountInWithFee = amountIn * 997;
amountOut = amountInWithFee * reserveOut / (reserveIn * 1000 + amountInWithFee);
}
function getAmountIn(uint256 amountOut, uint256 reserveIn, uint256 reserveOut) public pure returns (uint256 amountIn) {
require(amountOut > 0, "Router: insufficient output amount");
require(reserveIn > 0 && reserveOut > 0, "Router: insufficient liquidity");
amountIn = (reserveIn * amountOut * 1000) / ((reserveOut - amountOut) * 997) + 1;
}
// ---- views ----
function pairFor(address tokenA, address tokenB) public view returns (address pair) {
pair = factory.getPair(tokenA, tokenB);
require(pair != address(0), "Router: no pair");
}
function getReserves(address tokenA, address tokenB) public view returns (uint256 reserveA, uint256 reserveB) {
(address token0,) = sortTokens(tokenA, tokenB);
(uint112 r0, uint112 r1,) = IgneumPair(pairFor(tokenA, tokenB)).getReserves();
(reserveA, reserveB) = tokenA == token0 ? (r0, r1) : (r1, r0);
}
function getAmountsOut(uint256 amountIn, address[] memory path) public view returns (uint256[] memory amounts) {
require(path.length >= 2, "Router: invalid path");
amounts = new uint256[](path.length);
amounts[0] = amountIn;
for (uint256 i; i < path.length - 1; i++) {
(uint256 reserveIn, uint256 reserveOut) = getReserves(path[i], path[i + 1]);
amounts[i + 1] = getAmountOut(amounts[i], reserveIn, reserveOut);
}
}
function getAmountsIn(uint256 amountOut, address[] memory path) public view returns (uint256[] memory amounts) {
require(path.length >= 2, "Router: invalid path");
amounts = new uint256[](path.length);
amounts[amounts.length - 1] = amountOut;
for (uint256 i = path.length - 1; i > 0; i--) {
(uint256 reserveIn, uint256 reserveOut) = getReserves(path[i - 1], path[i]);
amounts[i - 1] = getAmountIn(amounts[i], reserveIn, reserveOut);
}
}
// ---- liquidity ----
function _addLiquidity(address tokenA, address tokenB, uint256 amountADesired, uint256 amountBDesired, uint256 amountAMin, uint256 amountBMin)
private
returns (uint256 amountA, uint256 amountB)
{
if (factory.getPair(tokenA, tokenB) == address(0)) {
factory.createPair(tokenA, tokenB);
}
(uint256 reserveA, uint256 reserveB) = getReserves(tokenA, tokenB);
if (reserveA == 0 && reserveB == 0) {
(amountA, amountB) = (amountADesired, amountBDesired);
} else {
uint256 amountBOptimal = quote(amountADesired, reserveA, reserveB);
if (amountBOptimal <= amountBDesired) {
require(amountBOptimal >= amountBMin, "Router: insufficient B amount");
(amountA, amountB) = (amountADesired, amountBOptimal);
} else {
uint256 amountAOptimal = quote(amountBDesired, reserveB, reserveA);
require(amountAOptimal <= amountADesired && amountAOptimal >= amountAMin, "Router: insufficient A amount");
(amountA, amountB) = (amountAOptimal, amountBDesired);
}
}
}
function addLiquidity(
address tokenA,
address tokenB,
uint256 amountADesired,
uint256 amountBDesired,
uint256 amountAMin,
uint256 amountBMin,
address to,
uint256 deadline
) external ensure(deadline) returns (uint256 amountA, uint256 amountB, uint256 liquidity) {
(amountA, amountB) = _addLiquidity(tokenA, tokenB, amountADesired, amountBDesired, amountAMin, amountBMin);
address pair = pairFor(tokenA, tokenB);
_pull(tokenA, msg.sender, pair, amountA);
_pull(tokenB, msg.sender, pair, amountB);
liquidity = IgneumPair(pair).mint(to);
}
function addLiquidityIGN(address token, uint256 amountTokenDesired, uint256 amountTokenMin, uint256 amountIGNMin, address to, uint256 deadline)
external
payable
ensure(deadline)
returns (uint256 amountToken, uint256 amountIGN, uint256 liquidity)
{
(amountToken, amountIGN) = _addLiquidity(token, WIGN, amountTokenDesired, msg.value, amountTokenMin, amountIGNMin);
address pair = pairFor(token, WIGN);
_pull(token, msg.sender, pair, amountToken);
IWIGN(WIGN).deposit{value: amountIGN}();
require(IWIGN(WIGN).transfer(pair, amountIGN), "Router: WIGN transfer failed");
liquidity = IgneumPair(pair).mint(to);
if (msg.value > amountIGN) _sendIGN(msg.sender, msg.value - amountIGN);
}
function removeLiquidity(address tokenA, address tokenB, uint256 liquidity, uint256 amountAMin, uint256 amountBMin, address to, uint256 deadline)
public
ensure(deadline)
returns (uint256 amountA, uint256 amountB)
{
address pair = pairFor(tokenA, tokenB);
_pull(pair, msg.sender, pair, liquidity);
(uint256 amount0, uint256 amount1) = IgneumPair(pair).burn(to);
(address token0,) = sortTokens(tokenA, tokenB);
(amountA, amountB) = tokenA == token0 ? (amount0, amount1) : (amount1, amount0);
require(amountA >= amountAMin, "Router: insufficient A amount");
require(amountB >= amountBMin, "Router: insufficient B amount");
}
function removeLiquidityIGN(address token, uint256 liquidity, uint256 amountTokenMin, uint256 amountIGNMin, address to, uint256 deadline)
external
ensure(deadline)
returns (uint256 amountToken, uint256 amountIGN)
{
(amountToken, amountIGN) = removeLiquidity(token, WIGN, liquidity, amountTokenMin, amountIGNMin, address(this), deadline);
require(IERC20Router(token).transfer(to, amountToken), "Router: token transfer failed");
IWIGN(WIGN).withdraw(amountIGN);
_sendIGN(to, amountIGN);
}
// ---- swaps ----
function _swap(uint256[] memory amounts, address[] memory path, address to_) private {
for (uint256 i; i < path.length - 1; i++) {
(address input, address output) = (path[i], path[i + 1]);
(address token0,) = sortTokens(input, output);
uint256 amountOut = amounts[i + 1];
(uint256 amount0Out, uint256 amount1Out) = input == token0 ? (uint256(0), amountOut) : (amountOut, uint256(0));
address to = i < path.length - 2 ? pairFor(output, path[i + 2]) : to_;
IgneumPair(pairFor(input, output)).swap(amount0Out, amount1Out, to);
}
}
function swapExactTokensForTokens(uint256 amountIn, uint256 amountOutMin, address[] calldata path, address to, uint256 deadline)
external
ensure(deadline)
returns (uint256[] memory amounts)
{
amounts = getAmountsOut(amountIn, path);
require(amounts[amounts.length - 1] >= amountOutMin, "Router: insufficient output amount");
_pull(path[0], msg.sender, pairFor(path[0], path[1]), amounts[0]);
_swap(amounts, path, to);
}
function swapTokensForExactTokens(uint256 amountOut, uint256 amountInMax, address[] calldata path, address to, uint256 deadline)
external
ensure(deadline)
returns (uint256[] memory amounts)
{
amounts = getAmountsIn(amountOut, path);
require(amounts[0] <= amountInMax, "Router: excessive input amount");
_pull(path[0], msg.sender, pairFor(path[0], path[1]), amounts[0]);
_swap(amounts, path, to);
}
function swapExactIGNForTokens(uint256 amountOutMin, address[] calldata path, address to, uint256 deadline)
external
payable
ensure(deadline)
returns (uint256[] memory amounts)
{
require(path[0] == WIGN, "Router: invalid path");
amounts = getAmountsOut(msg.value, path);
require(amounts[amounts.length - 1] >= amountOutMin, "Router: insufficient output amount");
IWIGN(WIGN).deposit{value: amounts[0]}();
require(IWIGN(WIGN).transfer(pairFor(path[0], path[1]), amounts[0]), "Router: WIGN transfer failed");
_swap(amounts, path, to);
}
function swapExactTokensForIGN(uint256 amountIn, uint256 amountOutMin, address[] calldata path, address to, uint256 deadline)
external
ensure(deadline)
returns (uint256[] memory amounts)
{
require(path[path.length - 1] == WIGN, "Router: invalid path");
amounts = getAmountsOut(amountIn, path);
require(amounts[amounts.length - 1] >= amountOutMin, "Router: insufficient output amount");
_pull(path[0], msg.sender, pairFor(path[0], path[1]), amounts[0]);
_swap(amounts, path, address(this));
IWIGN(WIGN).withdraw(amounts[amounts.length - 1]);
_sendIGN(to, amounts[amounts.length - 1]);
}
// ---- transfers ----
function _pull(address token, address from, address to, uint256 value) private {
(bool ok, bytes memory data) = token.call(abi.encodeWithSelector(IERC20Router.transferFrom.selector, from, to, value));
require(ok && (data.length == 0 || abi.decode(data, (bool))), "Router: transferFrom failed");
}
function _sendIGN(address to, uint256 value) private {
(bool ok,) = to.call{value: value}("");
require(ok, "Router: IGN send failed");
}
}

View file

@ -0,0 +1,32 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {ERC20} from "./ERC20.sol";
/// A test token with its own faucet: anyone takes DRIP tokens once an hour. Devnet 3, test tokens, no value.
contract TestToken is ERC20 {
uint256 public constant DRIP = 1_000 ether;
uint256 public constant DRIP_INTERVAL = 1 hours;
mapping(address => uint256) public lastDrip;
event Drip(address indexed to, uint256 value);
constructor(string memory name_, string memory symbol_) ERC20(name_, symbol_) {}
/// Mints DRIP tokens to the caller; refused inside DRIP_INTERVAL of the caller's last drip.
function drip() external {
uint256 last = lastDrip[msg.sender];
require(last == 0 || block.timestamp >= last + DRIP_INTERVAL, "TestToken: one drip an hour");
lastDrip[msg.sender] = block.timestamp;
_mint(msg.sender, DRIP);
emit Drip(msg.sender, DRIP);
}
/// Seconds until the caller may drip again (0 when it may).
function dripWait(address who) external view returns (uint256) {
uint256 last = lastDrip[who];
if (last == 0) return 0;
uint256 next = last + DRIP_INTERVAL;
return block.timestamp >= next ? 0 : next - block.timestamp;
}
}

View file

@ -0,0 +1,27 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {ERC20} from "./ERC20.sol";
/// Wrapped IGN, the WETH9 shape: one WIGN per IGN held by this contract, minted on deposit and burned on withdraw.
/// Devnet 3, test tokens, no value.
contract WIGN is ERC20("Wrapped IGN", "WIGN") {
event Deposit(address indexed to, uint256 value);
event Withdrawal(address indexed from, uint256 value);
receive() external payable {
deposit();
}
function deposit() public payable {
_mint(msg.sender, msg.value);
emit Deposit(msg.sender, msg.value);
}
function withdraw(uint256 value) external {
_burn(msg.sender, value);
emit Withdrawal(msg.sender, value);
(bool ok,) = msg.sender.call{value: value}("");
require(ok, "WIGN: send failed");
}
}

View file

@ -0,0 +1,137 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
import {Vm, VM_ADDRESS} from "./Vm.sol";
import {WIGN} from "../src/WIGN.sol";
import {TestToken} from "../src/TestToken.sol";
import {IgneumFactory, IgneumPair} from "../src/IgneumPair.sol";
import {IgneumRouter} from "../src/IgneumRouter.sol";
/// The AMM end to end on Foundry's EVM: pair creation, liquidity in and out, token and IGN swaps, the faucet's
/// hour, the k check. Every number is a test number: Devnet 3, test tokens, no value.
contract DexTest {
Vm constant vm = Vm(VM_ADDRESS);
WIGN wign;
TestToken tta;
TestToken ttb;
IgneumFactory factory;
IgneumRouter router;
address alice = address(0xA11CE);
address bob = address(0xB0B);
receive() external payable {}
function setUp() public {
wign = new WIGN();
tta = new TestToken("Test Token A", "TTA");
ttb = new TestToken("Test Token B", "TTB");
factory = new IgneumFactory();
router = new IgneumRouter(address(factory), address(wign));
vm.deal(alice, 1_000 ether);
vm.deal(bob, 1_000 ether);
}
function _drip(address who) internal {
vm.startPrank(who);
tta.drip();
ttb.drip();
tta.approve(address(router), type(uint256).max);
ttb.approve(address(router), type(uint256).max);
vm.stopPrank();
}
function _seed() internal returns (address pairAB, address pairAW) {
_drip(alice);
vm.startPrank(alice);
router.addLiquidity(address(tta), address(ttb), 500 ether, 500 ether, 0, 0, alice, block.timestamp + 60);
router.addLiquidityIGN{value: 100 ether}(address(tta), 500 ether, 0, 0, alice, block.timestamp + 60);
vm.stopPrank();
pairAB = factory.getPair(address(tta), address(ttb));
pairAW = factory.getPair(address(tta), address(wign));
}
function test_faucet_drips_once_an_hour() public {
vm.prank(bob);
tta.drip();
require(tta.balanceOf(bob) == 1_000 ether, "drip amount");
vm.prank(bob);
vm.expectRevert(bytes("TestToken: one drip an hour"));
tta.drip();
vm.warp(block.timestamp + 1 hours);
vm.prank(bob);
tta.drip();
require(tta.balanceOf(bob) == 2_000 ether, "second drip");
}
function test_add_liquidity_creates_pairs_and_mints() public {
(address pairAB, address pairAW) = _seed();
require(pairAB != address(0) && pairAW != address(0) && pairAB != pairAW, "pairs");
require(factory.allPairsLength() == 2, "two pairs");
(uint256 rA, uint256 rB) = router.getReserves(address(tta), address(ttb));
require(rA == 500 ether && rB == 500 ether, "AB reserves");
(uint256 rT, uint256 rW) = router.getReserves(address(tta), address(wign));
require(rT == 500 ether && rW == 100 ether, "AW reserves");
require(IgneumPair(pairAB).balanceOf(alice) == 500 ether - 1000, "LP minus the locked minimum");
require(IgneumPair(pairAB).balanceOf(address(0xdead)) == 1000, "locked minimum");
require(wign.balanceOf(pairAW) == 100 ether, "WIGN held by the pair");
}
function test_swap_exact_tokens_for_tokens_keeps_k() public {
(address pairAB,) = _seed();
_drip(bob);
address[] memory path = new address[](2);
path[0] = address(tta);
path[1] = address(ttb);
uint256[] memory quoted = router.getAmountsOut(10 ether, path);
// 10 in at 0.3 percent on 500/500: 10*997*500 / (500*1000 + 10*997) tokens
require(quoted[1] == 9775084808910328058, "quote");
(uint112 r0b, uint112 r1b,) = IgneumPair(pairAB).getReserves();
vm.prank(bob);
uint256[] memory amounts = router.swapExactTokensForTokens(10 ether, quoted[1], path, bob, block.timestamp + 60);
require(amounts[1] == quoted[1], "swap matches the quote");
require(ttb.balanceOf(bob) == 1_000 ether + quoted[1], "bob received");
(uint112 r0a, uint112 r1a,) = IgneumPair(pairAB).getReserves();
require(uint256(r0a) * uint256(r1a) >= uint256(r0b) * uint256(r1b), "k did not fall");
}
function test_swap_ign_both_ways() public {
_seed();
address[] memory path = new address[](2);
path[0] = address(wign);
path[1] = address(tta);
uint256 before = tta.balanceOf(bob);
vm.prank(bob);
uint256[] memory amounts = router.swapExactIGNForTokens{value: 1 ether}(0, path, bob, block.timestamp + 60);
require(tta.balanceOf(bob) == before + amounts[1] && amounts[1] > 4.9 ether && amounts[1] < 5 ether, "IGN to TTA");
path[0] = address(tta);
path[1] = address(wign);
uint256 ignBefore = bob.balance;
vm.startPrank(bob);
tta.approve(address(router), type(uint256).max);
uint256[] memory back = router.swapExactTokensForIGN(amounts[1], 0, path, bob, block.timestamp + 60);
vm.stopPrank();
require(bob.balance == ignBefore + back[1] && back[1] < 1 ether && back[1] > 0.99 ether, "TTA to IGN");
}
function test_remove_liquidity_returns_both_tokens() public {
(address pairAB,) = _seed();
uint256 lp = IgneumPair(pairAB).balanceOf(alice);
vm.startPrank(alice);
IgneumPair(pairAB).approve(address(router), lp);
(uint256 a, uint256 b) = router.removeLiquidity(address(tta), address(ttb), lp, 0, 0, alice, block.timestamp + 60);
vm.stopPrank();
require(a == 500 ether - 1000 && b == 500 ether - 1000, "pro rata minus the locked share");
require(IgneumPair(pairAB).balanceOf(alice) == 0, "LP burned");
}
function test_expired_deadline_is_refused() public {
_seed();
address[] memory path = new address[](2);
path[0] = address(tta);
path[1] = address(ttb);
vm.prank(alice);
vm.expectRevert(bytes("Router: expired"));
router.swapExactTokensForTokens(1 ether, 0, path, alice, block.timestamp - 1);
}
}

23
contracts/dex/test/Vm.sol Normal file
View file

@ -0,0 +1,23 @@
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
/// The Foundry cheatcodes this project uses, declared here so the tree needs no remote dependency.
interface Vm {
function startBroadcast(uint256 privateKey) external;
function stopBroadcast() external;
function envUint(string calldata name) external view returns (uint256);
function envAddress(string calldata name) external view returns (address);
function envOr(string calldata name, string calldata defaultValue) external view returns (string memory);
function toString(address value) external pure returns (string memory);
function toString(uint256 value) external pure returns (string memory);
function writeFile(string calldata path, string calldata data) external;
function deal(address who, uint256 newBalance) external;
function prank(address msgSender) external;
function startPrank(address msgSender) external;
function stopPrank() external;
function warp(uint256 newTimestamp) external;
function expectRevert(bytes calldata revertData) external;
function addr(uint256 privateKey) external pure returns (address);
}
address constant VM_ADDRESS = address(uint160(uint256(keccak256("hevm cheat code"))));

View file

@ -1,6 +1,6 @@
# Proving on AMD and Apple cards: what exists, what the CPU can do, what to tell the public
5 October 2026, from the project lead's two questions that evening: "test proving on the amd card?" and "can we test proving on
5 October 2026, from the founder's two questions that evening: "test proving on the amd card?" and "can we test proving on
mac?". PC 1 holds an RTX 5090 and an RX 9070 XT (gfx1201, 16 GB) in an eGPU; this Mac is an M5 Max. The prover is
SP1 (`proving/igneum-prove`, `docs/plans/proving-v0.md`, `proving-v1.md`), run on the GPU only through SP1's CUDA
server. Every figure below is measured (with its bench-log entry or job id) or cited (with its file or page); the
@ -115,7 +115,7 @@ gates or a backend ships; it is reviewed with every prover release.
| Consequence | Action | Owner |
|---|---|---|
| An AMD-only miner loses the proving share | the CPU tier of 4a is measured here (section 2); whether it becomes a tier is a decision for the project lead on those numbers | this analysis; the project lead |
| An AMD-only miner loses the proving share | the CPU tier of 4a is measured here (section 2); whether it becomes a tier is a decision for the founder on those numbers | this analysis; the founder |
| The rig's prover unit must select NVIDIA cards only | already true in `prover_decision`; told the rig-installer agent to keep it as a stated rule and to print the CPU-fallback line for AMD-only rigs | rig-installer agent |
| The app's Proving tile on an AMD-only or Apple machine should say why it is off and name the CPU path | the `provedefault.rs` lines already say so for Apple; AMD-only Windows machines get "no NVIDIA card ..." | proving agent (told) |
| The site and litepaper over-promise for AMD and Apple | the line of 4c, to land with the next site pass (copy law; `node site/build.mjs`; link-check) | site-pages owner; not changed here |

View file

@ -1,8 +1,8 @@
# ASIC resistance, 2011 to 2026: the history, the papers, the lessons, and the audit of Igneum against them
5 October 2026 (night), branch `asic-history`. Asked by the project lead at 20:05 UTC: "do a full on deep dive into the full history of 'asic resistance' and see if we can add or upgrade anything." Baseline for the audit: the Counter ASIC 2.0 final class decided tonight (`docs/plans/counter-asic-2-status.md` on `ca2-coord`, entries 20:16 to 22:25 UTC; `docs/analysis/chip-model-v3.md` on `ca2-mixer` 1ab8b21). Every figure about another chain cites a repo file, a paper or a dated article, or is labelled approximate. Hash-per-joule gains are computed from the cited hashrate and watt figures of the chip and of the best consumer GPU of the same year, and are approximate by construction (GPU figures vary by tuning). Research gathered by four sub-agents between 20:10 and 20:45 UTC; the fetch failures they reported are listed in section 6.
5 October 2026 (night), branch `asic-history`. Asked by the founder at 20:05 UTC: "do a full on deep dive into the full history of 'asic resistance' and see if we can add or upgrade anything." Baseline for the audit: the Counter ASIC 2.0 final class decided tonight (`docs/plans/counter-asic-2-status.md` on `ca2-coord`, entries 20:16 to 22:25 UTC; `docs/analysis/chip-model-v3.md` on `ca2-mixer` 1ab8b21). Every figure about another chain cites a repo file, a paper or a dated article, or is labelled approximate. Hash-per-joule gains are computed from the cited hashrate and watt figures of the chip and of the best consumer GPU of the same year, and are approximate by construction (GPU figures vary by tuning). Research gathered by four sub-agents between 20:10 and 20:45 UTC; the fetch failures they reported are listed in section 6.
## 0. One page for the project lead
## 0. One page for the founder
**What the history says Igneum is doing right.**
@ -228,7 +228,7 @@ Ranked by how much the history says each would change the outcome, with the cost
| 1 | **Price the partial-store chip and draw the time-memory curve.** A chip that stores a fraction f of the dataset in HBM or on many narrow DRAM channels, recomputes the rest from a 256 MiB on-die cache under x8, and reads with 4-byte granularity. Rows for f = 0.25, 0.5, 1 at HBM3 and at GDDR7 random-read rates, priced in energy per hash (Rao's metric) and in reads in flight per watt | The only chip class that beat a memory-bound GPU hash: Ethash's 2.1x to 4.8x came from the memory system with no on-die dataset (rows 3, 4); Rao priced a 16-die DAG holder under a GPU board in 2019; Cuckoo's curve was wrong by 50x until drawn [P20]; O-1.6 is open and `MEMHARD.md` section 3 item 2 says the curve was never drawn | None (analysis) | The row may come out over 2x, which would qualify the public claim before anyone else does | Genesis (before the vectors freeze) |
| 2 | **A random item-derivation program per day** in place of the fixed-shape mixer: a SuperscalarHash-style generator, integer only, drawn from the day key, with its own acceptance test, compiled once a day by miners and verifiers | RandomX's reason for SuperscalarHash: a fixed derivation is hard-wired by a chip; a random one makes the light-mode chip a CPU (section 2.4). In Igneum's model the fixed shape is the 3x factor that turns 0.31x into 0.92x; removing the factor is worth more than x8 to x16 would be (x16: 0.46x with the factor by M16's table) | None per hash (the daily build is 23 to 77 ms at x8 and would roughly double); the verifier needs a per-day compiled derivation (a JIT, or a round schedule drawn from a fixed set of reviewed rounds), measured against the 10 ms gate | Cryptanalysis of random ARX programs; weak draws; a JIT in the verifier is new attack surface; the vendors must agree bit-exactly on a program they compile | Reserve (named family, unlock by height or signal) now; genesis if the verifier cost is measured under the gate before the freeze |
| 3 | **External cryptanalysis of M_r, the chained cache and the acceptance rule before genesis**, with the x8 shape as the target | Lesson 9 (MTP, Catena, Argon2i, Cuckoo); RandomX bought four audits for $141,000 before launch [S60]; the x8 decision multiplies the mixer's weight in the chip model, so a shortcut inside the mixer is now worth 8x more to a chip | None | Finding something late moves the vectors; not finding it in time moves nothing | Genesis gate (ledger M7, raised in priority) |
| 4 | **The clock and the detector.** (a) A share-pattern detector on the observer: per-program hash-rate spread, nonce-group patterns and per-card-model rate bands, with an alert when a population behaves like one fixed design (MoneroCrusher's method); (b) a stated trigger: the bounty escrowed and the benchmark live before daily issuance crosses about $50K (Vorick's rule), not on a calendar date | Lesson 10 (85% secret share); section 2.5's table (chips at $20K to $30K a day on compute-bound hashes); D11 (the bounty is unfunded) | None | A detector with false positives; a trigger the project lead has to fund | Not a layer; genesis-independent; do it before the public testnet |
| 4 | **The clock and the detector.** (a) A share-pattern detector on the observer: per-program hash-rate spread, nonce-group patterns and per-card-model rate bands, with an alert when a population behaves like one fixed design (MoneroCrusher's method); (b) a stated trigger: the bounty escrowed and the benchmark live before daily issuance crosses about $50K (Vorick's rule), not on a calendar date | Lesson 10 (85% secret share); section 2.5's table (chips at $20K to $30K a day on compute-bound hashes); D11 (the bounty is unfunded) | None | A detector with false positives; a trigger the founder has to fund | Not a layer; genesis-independent; do it before the public testnet |
| 5 | **Rank layer 9 (the epoch length) up, and measure the FPGA lane**: the compile-ahead cost per card at a 10-minute epoch (the `ca2-epoch` work), plus an estimate of a soft-overlay FPGA miner with HBM (reads in flight per watt against the 5090's 17.5 G/s) | FPGAs were the first adversary of Lyra2REv2 and X16R and came back within weeks of X16Rv2 (rows 7, 9); Xelis forked for FPGA resistance (row 30); a per-hour program is a bitstream target in a way a per-hash program is not | At 10-minute epochs: 6x the compile work per card (measured on `ca2-epoch`); the VDF lead shrinks | A short epoch moves the difficulty window (spec 1.12) and the seed path | Reserve (as decided), with the measurement before the public testnet |
| 6 | **Order the reserve by chip-unfriendliness**: families that force a full 32-bit datapath per lane first (byte permute, bit-field extract, variable shifts, popcount, select, the second shuffle form), mm8 last | Least Authority's "watch ML hardware"; int8 matrix blocks are licensable IP at every node; Apple pays 1.6x to 4.7x per emulated dot4 (status 20:38) | None at launch | None | Reserve ordering, genesis |
| 7 | **A vendor-share metric and a 3.0 target for the AMD gap**: the share of hashrate by vendor published with the benchmark, and the line-width question kept open as the plan says | Lesson 8: a one-vendor fleet is a softer version of chip capture; Equihash's NVIDIA tilt and Ethash's balance were part of each chain's miner politics (rows 3, 5) | n/a | A width that closes the gap makes the 5090 bandwidth-bound (status 20:27) | Counter ASIC 3.0 |
@ -256,14 +256,14 @@ Evaluated and placed nowhere, with the reason:
| Divergent data-dependent branches | Nowhere (already excluded) | Branches cost a GPU divergence and a chip nothing; RandomX's single predictable branch targets speculative CPUs, which Igneum does not have |
| Floating point | Nowhere (already excluded) | Vendor rounding splits the chain (spec 1.14); RandomX could afford it because its target is one ISA family with IEEE semantics |
## 5. Decisions this raises for the project lead
## 5. Decisions this raises for the founder
| # | Decision | Recommendation |
|---|---|---|
| 1 | Add the partial-store chip rows to `chip-model-v3.md` and draw the time-memory curve before the public testnet | Yes, before the vectors freeze (addition 1) |
| 2 | Name a random item-derivation program as a reserve family, and fund the verifier measurement that would move it to genesis | Reserve now; genesis if the verifier lands under the gate (addition 2) |
| 3 | Commission the external cryptanalysis of M_r and the chained cache before genesis, with the x8 shape as the target | Yes (addition 3; ledger M7) |
| 4 | Escrow the bounty and set its trigger to daily issuance, not to a date; build the share-pattern detector on the observer | Yes to the detector now; the escrow is the project lead's (D11) |
| 4 | Escrow the bounty and set its trigger to daily issuance, not to a date; build the share-pattern detector on the observer | Yes to the detector now; the escrow is the founder's (D11) |
| 5 | Rank the epoch-length reserve above the mm8 reserve, and measure the FPGA lane | Yes (additions 5 and 6) |
## 6. Sources and limits of this research

View file

@ -0,0 +1,635 @@
# Internal attack pass before the freeze (F1 to F10)
The internal cryptanalysis pass of `docs/plans/cryptanalysis.md` section 4.2, run before the freeze tag
`cryptanalysis-target-1`, so the adversarial lanes confirm rather than discover. The founder's word, 7 October 2026:
"make sure they find ZERO flaws". Every finding is ours, fixed and re-gated, before the public report.
Target: the hash class the chain runs after 0.3.15's flip, `igneum-pow` generator v4 `V4_CLASS` = `mx8+sh256x27`
(`LoadClass::MX8`, `ShadowClass { instrs: 256, reps: 27 }`), the acceptance rule, the verifier, the era draw,
the latency-shadow dataset and ladder, the chip and FPGA cost model. Scope and gates are section 1.1 and 1.4 of
the plan, the same tests the public report is held to.
Lane: attack-pass, worktree `igneum-wt-attack`, branch `attack-pass` from `origin/master` `ab99e5e3`.
Binary built on igneum-build-1 (ELF x86-64, `igneum-pow` 0.2.0, sha256 6d2867...1a9ebe5) and run there under
the box's slots; model and era work from `sim/horizon/algorithm/model.py` and `infra/fast-time/`. Each row below
carries the method, the known-failed shape where one exists, the result with numbers, and PASS, RUNNING,
BLOCKED or FINDING. PASS RECORD (7 October 2026, 16:0x UTC, 17:0x UK): every row reads PASS or FIXED-AND-PASSED. F1 PASS (AP-F1-1 on the
v5 list at 3.0 percent); F2 PASS, effort-bounded; F3 PASS; F4 PASS against class v4 (AP-F4-1 on the v5 list); F5
FIXED-AND-PASSED (the F2 hour skipped by decision); F6 PASS (the worst of 10^5 programs 8.708 ms on the half-core
proxy; O-1.14 closed on an i7-9700K); F7 PASS on all three sub-rows (the era VDF in the node, 0 of 6 re-rolls); F8
FIXED-AND-PASSED in class v4 sub-version 3 at 017e7037 (AP-F8-1, AP-F8-2, AP-F8-3 ours and closed; the four-seed
tail attributed to per-site bucket concentration, 8 October 2026); F9 PASS on (a) and (c), (b) closed by the same fix; F10 PASS. Frozen generator: class v4
sub-version 3, igneum-pow 017e70376489251e18564c0abce7e466e606c8b3, devnet epoch-0 id a785001687d8688a. The
freeze tag `cryptanalysis-target-1` is the coordinator's cut on this record; the era VDF precondition is met (F7 a). Main checks every number
against the log before quoting it to the founder.
## Status board
| # | Attack | Gate (same as 1.4) | Result so far | Status |
|---|---|---|---|---|
| F1 | Shadow block compressibility and shortcut search | no compression of the shadow block beyond the honest compiler's simplification, measured against that compiler on the same program; the 27 repetitions never fewer than 27x (coordinator's ruling 7 Oct 2026, 12:3x UK; the plan's gate (1) and row F1 carry the same) | 10^4 and 10^5 class v4 programs: saved instructions mean 0.62%, max 5.078% at 10^5 (1 of 100,000 over 5%, 0 over 10%); nothing folds or dedupes across the 27 passes; every saving is local peephole algebra that clang -O3 removes from the honest kernel too (IR counts match on the worst programs), so against the compiler the compression is 0; 0 mismatches in 2 x 10^5 differential and verifier checks, z3 window proofs 0 counterexamples. AP-F1-1 routed to the v5 list as a shadow redundancy bound. Record `docs/analysis/attack-pass/f1-shadow.md` | PASS; AP-F1-1 on the v5 list |
| F2 | Mixer round margin (SAT/MILP, 1 to 4 keyed applications) | no distinguisher or shortcut beyond 2 of the 8 applications | one application characterised (differential weight 10 to 12, linear 1, verified on the real code on three days); two applications: no trail at or below weight 20 to 24 within 7,200 s per job, the MSB and LSB families die at two; rotational-XOR no bias at one application; the multiply layer folds on 0 of 2^20 inputs, k applications cost k; three applications: no differential trail at or below weight 29 to 35 and no linear at or below 24 to 28, four: 39 to 47 and 24, every job at its 7,200 s cap. Record `docs/analysis/attack-pass/f2-mixer.md` | PASS (effort-bounded) |
| F3 | Chained cache j+1 bound and storage-vs-recompute curve | no derivation under j+1 blocks; curve monotone; f=1 point unchanged | 0 of 64 and 0 of 1,024 lines under j+1 (exhaustive closure search, cross-checked by exhaustive pebbling at 10 lines, 10,240 pairs, 0 mismatches); both planted broken chains fire; curve monotone at both op counts; f=1 point 9,360 ops per item unchanged. Record `docs/analysis/attack-pass/f3-cache.md` | PASS |
| F4 | Weak-day census over 2^24 day keys | fraction of days with gain over 1.1x under 2^-20 | PASS against M2 (DSP-bound datapath): 0 of 2^28 days over 1.1x; planted weak days fire; every ROT and RC class 0. Bound finding AP-F4-1 on M1 (LUT adders): 5,476 of 2^24 days (3.26e-4) over 1.1x as the tail of a sum, no weak class; worst public-calendar day 29,337 at 1.121x, at most 12.1% more rate that day for a per-day LUT FPGA, 0 for any chip; redraw rule (NAF sum under 163 rejected) routed to the next class. Record `docs/analysis/attack-pass/f4-weakday.md` | PASS (v4); AP-F4-1 routed to the next class |
| F5 | Chip-model sweep + AWS F2 FPGA hour | evidence row 17 holds across the sweep; FPGA row under 27 M reads/s/W | sweep: 2.1x at k=1 GDDR7 reproduces, 3.2x at k=0.5, 4.1x at k=0.3 (matches ledger M32); FPGA row 2.3 to 2.9 G/s, 10 to 20 M reads/s/W (literature). FINDING: the k=0.33 figure is framed as the X9's measured core (M32) and a "measured class" (ladder branch §5a); the X9 was withdrawn before launch and never benchmarked. F2 hour SKIPPED: no AWS account | FIXED-AND-PASSED (sweep PASS; AP-F5-1 fixed and re-gated 7 Oct 2026: chip section re-run 2.1x at k=1 unchanged, identity grep 0 hits, site lane concurred); F2 hour SKIPPED-BY-DECISION (the founder, 7 Oct 2026, 09:5x UK; plan 4.2 row F5 is the sweep only at 3714c2a0; the FPGA row stays the JEDEC-ceiling model row labelled unmeasured) |
| F6 | Verifier worst case over 10^5 programs + O-1.14 laptop run | worst program under 10 ms cold on the half-core proxy and the laptop | 100,000 programs ranked, 50,000 timed cold one-core (worst 6.194 ms), the worst 200 re-timed and the worst 1,000 timed on the half-core proxy under the per-core lease (core 40 at 3,799.9 MHz): worst 8.708 ms (`attack-f6/87142`), 1.29 ms under the gate, every half-core reading under 9 ms; dr736 fails as it must (15.49). O-1.14 CLOSED on an i7-9700K (v4 6.334 ms cold max, dr736 10.04 fails). Ladder ceiling from the worst program on the half-core proxy: N about 300,000, so rung 2 admissible, rung 3 not. Record `docs/analysis/attack-pass/f6-verifier.md` | PASS |
| F7 | Era-draw bias harness + 2^20 era-seed census | no re-roll inside the publish window; no era class with gain over 1.1x over 2^-20 | (b) census PASS at 2^24 seeds (no class over 1.1x, stride bijective, R, pos and M uniform, planted cases fire); (c) 64-bit day-key seeding PASS (0 collisions in 2^17); (a) PASS with the era VDF in the node (era-vdf lane, fork era-vdf-node 394a5902 on release-0.3.20-node c4459193, behind era_vdf_activation_daa; master a4eaf766 carries the spec text, `docs/analysis/era-vdf-2026-10-07.md` and `tools/era-vdf/reroll.mjs`): against the real era cut on three fast-time nodes the re-roll harness fires with the VDF off (6 of 6 cuts) and is silent with it on (0 of 6; the adversary's 5.4 to 6.6 s evaluation against a 1 s block interval, the honest chain 3 to 10 blocks ahead, three nodes agreeing on every era seed); the delay is 517 s on the fastest prover measured (chiavdf NUDUPL over GMP, 208.8K squarings/s) at T = 108,000,000, 259x the 2 s window. Open, not a gate: the verify is 22 ms against the 10 ms target (O-4.6, 0.3.22). Record `docs/analysis/attack-pass/f7-era.md` | PASS (a, b, c) |
| F8 | Uniformity censuses (line-index 2^28, distinct lines, cross-hash histogram) | uniform within the window model of spec 1.13.1, layer 8; the excess beyond it within 6 sigma over 64 seeds; no hot set under 1% of items beyond the model | line index PASS at 2^28. The 6 Oct stream: 31 of 64 seeds over 1.2x (AP-F8-1, the lossy load source). Sub-version 1 (8c728ca3): 11 of 64. Sub-version 2 (07a809a7 / 8bdcbdd8): 9 of 64; AP-F8-2 (exhaustion) closed there, 0 of 10^6. Sub-version 3 (017e7037: the acceptance executing the shadow block, AP-F8-3; the shared-operand rule; (c'') the distinct-index ratio): 60 of 64 under 1.2x, the four over the named tail at 1.22x to 1.50x, attributed 8 October 2026 to a quarter-bit bucket concentration at one narrow-window site each, with no chip consequence; 0 exhausted in 24,631 chain-shaped seeds, the draw total by construction. Record `docs/analysis/attack-pass/f8-uniform.md` | FIXED-AND-PASSED (sub-version 3, 017e7037, the frozen generator) |
| F9 | Acceptance edges (39) + header grinding on an RTX 5090 | zero passing programs with a hot set under 1%; grinding gain under 1% of rate | (a) edges on generator 4 over 10^5 seeds: 34 disagreements in 105,064 candidates (29 const_bit, 4 bias, 1 lane_const), every one a one-bit or sampling-noise property moving the choice to an attempt both stand-ins accept, nothing in the attacker's favour: PASS; (b) hot-set search over 10^6 seeds: 11,696 passing programs (1.17%) concentrate 1% or more of reads on a hot set, worst 17.3%: FINDING, the same or-saturation load-source class as AP-F8-1 found by a second harness (F9-1 merged into AP-F8-1), re-gated on the amended stream; (c) grinding on the 5090: +0.004% at K = 2^14, ceiling +43%: PASS. Record `docs/analysis/attack-pass/f9-grind.md` | (a) PASS; (b) the AP-F8-1 class on the 6 Oct stream, closed by sub-version 3 (F8's census is the re-gate instrument; this harness's hot-share metric counts the era's windows); (c) PASS |
| F10 | Ladder signal monotonicity harness | no step without 90% over 7 windows in either direction | known-fail fails, known pass passes; new cases on the exact-share driver: 89% up holds (no step), 89 then 90% down with restarts steps only at 90% after the 7-window cool-down, the floor holds under 100% down (never below rung 0); decision per seed block memoised, identical after a restart, never differs between nodes; 19 of 19 checks per case. Two items to main, not findings: a stale commit string in the ladder lane's igneumd, and proof-synced nodes deciding rung 0 until the witness lands (a precondition line for spec 01). Record `docs/analysis/attack-pass/f10-ladder.md` | PASS |
## The rows
### F1. Shadow block compressibility and shortcut search (hash lane)
Method: over 10^4 class v4 programs, constant folding, dead-register elimination, common subexpressions across the
27 repetitions, linear sub-block detection, SAT equivalence on reduced blocks; the minimum op count per program
against N. Known-failed shape: a shadow that constant-folds or dedupes across its 27 identical passes so a chip
pays fewer than 55,296 shadow instructions per hash. Entry point: `igneum-pow show --program-class v4` prints the
256-instruction shadow (op mix add=47 rotl=30 xor=30 shfl=29 mad=27 mul=22 sub=21 rotr=20 mulhi=18 or=12 on the
genesis seed). Gate: best compressed block within 5% of N on every program; no program over 10% compressible.
Result: RUNNING. What a failure moves: an acceptance-rule line for the shadow block (rule (c) runs the block),
packs re-cut.
### F2. Mixer round margin (hash lane, on the box)
Method: SAT or MILP differential and linear search on 1 to 4 keyed applications with drawn rotations; rotational-XOR
on the ARX layer; the fold of the multiply layer across applications checked algebraically. Known-failed shape: a
differential or linear trail or an algebraic fold that distinguishes or shortcuts more than 2 of the 8 applications
between dependent reads. Gate: no distinguisher or shortcut beyond 2 of the 8 applications. Result: RUNNING (prior
`ca2-mixer` evidence to be re-gated). What a failure moves: `mixer_mult` 16 or a shape change; verifier re-measured.
### F3. Chained cache j+1 bound and storage-vs-recompute curve (hash lane, on the box)
Method: exhaustive search on a 2^10-line model segment for a line derivable without an earlier line; the curve from
f = 1/64 to 1 in ops per item. Known-failed shape: a line (s, j) computable in fewer than j+1 block evaluations
without an earlier line (the MTP address-steering break shape). Gate: no derivation under j+1 blocks; curve monotone;
f=1 point unchanged. Result, 7 October 2026, 09:10 to 09:12 UK on the box (`docs/analysis/attack-pass/f3-cache.md`; logs
`/srv/builds/igneum-wt-attack/attack-f3/r1-*.log`): PASS on all three clauses. The chain extracted from
`Cache::fill_segment` (verified equal to the code on 16 of 16 key and segment pairs; `block == chacha_block` on
100,000 random inputs) has line j fed by line j - 1 only; the exhaustive closure search finds 0 of 64 and 0 of
1,024 lines under j + 1 (every line costs exactly j + 1), cross-checked by an exhaustive pebbling search at 10
lines (10,240 configuration and target pairs, 0 mismatches). The two planted chains fire: `skip2` (63 of 64 under
j + 1) and `nofeed` (every line in 1 block). The curve over stored cache lines is monotone non-increasing from
f = 1/64 to 1 at 608 (counted) and 700 (MEMHARD.md) ops per block; the f = 1 point is 9,360 ops per item,
41.7 MH/s at the 50 T op/s budget, unchanged. `ca2-cache` was the hot-table experiment, not a chain analysis, so
there was nothing to re-gate. Observation (coordinator and the F3 record, not a finding): `funding.md` B2 rank 2
prices the trade-off at the naive placement; the optimal placement of every 8th line costs 3.17 blocks per read,
not 3.5, and 16.0 at f = 1/64, not 31.5 (brute force over 4,426,165,368 sets at n = 8); the chip stays worse than
the full mirror at every f under 1, so the verdict stands, and a `chacha_block` shortcut in chaining mode stays
the open question of the adversarial lanes. What a failure would have moved: the chain construction (a second feed-forward or
a cross-segment tie).
### F4. Weak-day census over 2^24 day keys (hash lane, on the box)
Method: 2^24 day keys through `MixParams::with_shape`; the ROT classes (all equal, complementary pairs, small
amounts), MUL low weight, RC structure, each per-day gain measured on the box verifier. Known-failed shape: a day
key whose drawn ROT/MUL/RC gives a fixed datapath a gain over 1.1x (the "weaker authorized parameters" class,
Kudelski 2019). Gate: the fraction of days with any gain over 1.1x under 2^-20. Result: RUNNING. What a failure
moves: a rejection-and-redraw rule on the draws.
### F5. Chip-model sweep and the FPGA hour (algorithm lane)
Method: `sim/horizon/algorithm/model.py` over k 0.2 to 1.5, tFAW 12 and 28 ns, HBM4 2.3 and 21.4 G reads per
stack, amortisation 1 to 3 years, electricity USD 0.05 to 0.15 per kWh; and the AWS F2 hour replacing the FPGA
ceiling row with a measurement. Known-failed shape: an input of the published model that, when corrected, lifts the
f=1 chip's per-joule edge over the 5090 above the published 2.1x at k=1.
Gate: the published sentence (evidence row 17) holds across the sweep; the FPGA row under 27 M reads/s/W.
Result (sweep): PASS on the numbers. The model's measured-anchor column (GDDR7, the 5090 reads 82% of its ceiling)
gives the f=1 chip's v4 per-joule edge over the RTX 5090 bench row as 4.1x / 3.2x / 2.1x / 1.5x at k = 0.3 / 0.5 /
1 / 1.5. At k = 1 the figure is 2.1x, and 3.9x at k about 0.33, which matches `fud-ledger.md` M32. The higher HBM3
and HBM4 columns rest on an 8-activate per 12 ns window that JEDEC HBM2 timings (4 per 28 ns) do not support; the
model already states GDDR7 is the column to quote. One wording gap: `evidence.md` row 17 says "brings it to about
2x", which is a floor that holds at k about 0.9 and above but understates the edge at lower k (3.2x at k = 0.5). The
accurate statement is M32's, 2.1x at k = 1 with the k range beside it. The sweep's numbers stand; the finding is the
X9 framing below.
FPGA row: the HBM2 FPGA ceiling is 2.3 to 2.9 G reads/s (measured Shuhai U280, FCCM 2020, equal to the JEDEC
tFAW-bound 2.3 G/s), 10 to 21 M reads/s/W at 115 to 150 W, 0.30 to 0.47x of the 5090 per watt. Under the 27 M
reads/s/W gate. The AWS F2 hour is SKIPPED-BY-DECISION (the founder, 7 October 2026, 09:5x UK: not needed for now, not blocked;
plan 4.2 row F5 at commit 3714c2a0 on branch cryptanalysis is the chip-model sweep only, 4 h, the algorithm lane).
There is also no AWS account or `aws` CLI on this Mac. The FPGA row stays the JEDEC-ceiling model row labelled
unmeasured; the chip-model lane prices it from the reads-in-flight model; the public report states that the F2 measurement was not run.
FINDING (X9 framing), owning lane algorithm and hash (the ladder lane is closed, so ours): the published numbers
already carry 2.1x at k = 1 beside 3.9x at k about 0.33 (`fud-ledger.md` M32, recalibrated under X35). The error is
the framing. M32 calls the k = 0.33 figure "the X9's core" and the `ladder` branch's `latency-ladder.md` section 5a
calls k about 0.33 a "measured class". Bitmain's Antminer X9 (RandomX ASIC, 1 MH/s, 2,472 W, about USD 5,600) was
announced and, per pcpraha.cz ("Antminer X9 canceled: Bitmain withdraws model from market before launch") and
r/MoneroMining, withdrawn before launch. Its implied core efficiency (k about 0.33) is a CLAIMED datasheet figure
from a design that never shipped and was never benchmarked, not a measured calibration point. It is carried as the
pessimistic bound, not a calibration. This collides with the merged ledger X34 ("RandomX has a shipping chip;
correct every sentence that said otherwise"): if the X9 was withdrawn, X34's correction is itself wrong and must be
reversed. Confirmed from primary sources (coordinator, 7 October 2026): pre-orders opened 26 December 2025 (shipments
scheduled for July 2026), withdrawn in mid-May 2026 with buyers refunded before any unit shipped, no independent
benchmark, Bitmain never published a cancellation (its shop lists it as sold out), and the box was commodity Sophgo
SG2044 server SoCs with an AES accelerator and 60-plus DRAM sticks, no tapeout; its claimed edge about 2x per joule
over a tuned Zen 4 part, about 3x over a stock desktop CPU. Re-cut applied the same day: `evidence.md` row 17 (claim
and measured cells) and `fud-ledger.md` M32's answer paragraph on attack-pass, and the ladder design doc section 5a
on branch `attack-ladder-5a` from the ladder tip 7003f9f5 (the `ladder` branch is checked out by another lane, so
the fix rides its own branch for the ladder owner to take). The served-text rows (X34 reversal, X36) belong to the
site lane, which confirmed the wording agrees. What a finding moves
(plan 4.2 F5): the sentence re-cut before the freeze so the public report carries the corrected model. The re-cut, once the
fact is confirmed: k about 0.33 labelled a claimed pessimistic bound from a withdrawn design everywhere it appears;
2.1x at k = 1 on the GDDR7 measured anchor kept as the headline with the k range beside it; k itself unmeasured
until the chip-model lane produces it. This row reads FIXED-AND-PASSED only after the re-cut and its re-gate.
Economic row the withdrawal implies: a recompute chip at a 3x fixed-function factor against a CPU and GPU fleet
must recover its NRE (low to mid seven figures at a modern node, `chip-model-v3.md`) and carry a fork threat (a
class change at 95% miner signal can redraw the datapath the chip bakes in). The X9 at 2.47 J per KH against a
RandomX CPU fleet did not clear that bar at Monero's hash and price; the same arithmetic against Igneum's class v4,
with the shadow block and the automatic era draw as extra firmware risk, is why the chip model's verdict is a
deliverable and not a courtesy (plan 2.3). The confirmed reading: a box with a 2x to 3x per-joule edge and no NRE (commodity SoCs) was withdrawn rather than
face a 1.5x re-tune of RandomX, so the tapeout economics of a 3x chip against Igneum are worse than the X9's. This
row is the pessimistic case, not a measured gain.
### F6. Verifier worst case (algorithm lane)
Method: 10^5 class v4 programs timed on the box one-core and half-core proxies for the slowest warp (base program
and shadow block), plus the O-1.14 laptop run (the Windows `igneum-pow` build on the box, the relay, `bench
--warps 50`). Known-failed shape: a drawn program whose verifier warp exceeds 10 ms cold (the acceptance rule bounds
the miner's side, not the verifier's; `dr736` already FAILs at 10.51 ms cold one-core, but it is not the shipping
class). Gate: the worst program under 10 ms cold on the half-core proxy and on the laptop.
O-1.14 route (coordinator, 7 October 2026, 10:1x UK): no US laptop is due, so the laptop run is replaced by a
rented 2019-class CPU host through the fleet agent, capped at two hours of rent; the Linux `igneum-pow` from the box
(the same binary as the proxies) runs `bench --warps 50` for v2, mx8, mx8+sh256x27, dr368 and dr736 (the
known-fail) on it; INCOMPLETE with the numbers so far if the cap lands first. The Windows exe was also built on the
box for the day a laptop appears (1,009,675 bytes, sha256 fbed7538...e6c9). Outcome, 09:40 UK: no 2019-class CPU
host stood up. Vast accepted and dropped five CPU-class rents within 30 s each (i7-9700K, i5-8500, Xeon W-2133 and
W-2123) under the account's automatic new-account spend limit (support ticket open since 6 October), and RunPod has
no 2019-class CPU pod; so O-1.14 reads INCOMPLETE on the box proxies today and stays a precondition of the freeze.
Next try: the US laptop when it registers on the relay (the exe is ready), or Vast once the spend limit lifts; the
bench script is staged and runs in minutes. Correction, 10:4x UK (fleet agent): the provider dropped nothing; every
rent stood up and ran, hidden by Vast's instance listing cap of 25 rows on an account holding 38, so the hosts sat
idle and were destroyed. The fallback is re-rented under the same word (i7-9700K class, two-hour cap from its start,
read by id); the result replaces this line when it lands.
O-1.14 RESULT, 7 October 2026, 09:49 UK, Vast instance 54613164, Intel Core i7-9700K (2019 desktop core, Coffee Lake,
read at 4,170 MHz during the run, 31 GB DDR4, Ubuntu 24.04), the box-built Linux `igneum-pow` (sha256 6d286783...),
`bench --seed igneum-genesis --day 2026-10-03 --warps 50` on one core (`taskset -c 1`), the host otherwise idle; log
`docs/analysis/attack-pass/o114-i7-9700K-2026-10-07.log`:
| Class | Cold max (ms per warp) | Average of 50 (ms) | Gate 10 ms | Box one-core cold | Box half-core |
|---|---|---|---|---|---|
| v2 | 1.582 | 1.280 | pass | 1.30 | n/a |
| mx8 (class v3) | 5.394 | 5.267 | pass | 4.67 | 7.56 |
| mx8+sh256x27 (class v4, the target) | 6.334 | 6.006 | pass, 3.7 ms of headroom | 5.06 | 8.23 |
| dr368 | 5.540 | 5.426 | pass | 5.32 | 8.16 |
| dr736 (the known-fail) | 10.290 | 10.042 | FAIL, as it must | 10.51 | 15.49 |
Cache fill 276 ms on the 9700K core (box 361 ms, M5 Max 175 to 181 ms). The 2019 desktop core sits between the box's
two proxies as the arithmetic predicted (1.2x the box one-core cold on v4, 0.77x the half-core); the known-fail
fires on it. A 2019 laptop core at 3.5 GHz reads about 15 to 20 percent slower than this desktop part (approximate,
clock ratio), so about 7.0 to 7.6 ms on v4, still under 10 ms. Consequences: a 2019-class node verifying class v4
spends 0.6 percent of one core at 1 bps and 6 percent at 10 bps; a header flood needs about 160 invalid headers a
second to saturate one such core; a pool verifies about 160 shares a second per core; IBD of 108,000 headers is
about 11 minutes of one core. The implied ladder ceiling on this core: the shadow costs 0.74 ms per 55,296
instructions (v4 minus mx8), so the 4.0 ms of headroom buys about 300,000 more shadow instructions, N about 650,000
counted ops at the 1.83 convention (approximate), against 370,000 on the half-core proxy and 1,060,000 on the
2.5x rule; the half-core proxy stays the standing pessimistic rule and the ladder's ceiling should be taken from
it, not from this desktop part. O-1.14 is CLOSED on a real 2019-class core for the genesis program; the F6 row
still owes the 10^5-program worst case before it reads PASS.
Result (average, verified on the box): class v4 `mx8+sh256x27` runs 4.90 to 5.06 ms per warp cold on one EPYC
9454P core (nice 19, taskset), 8.23 ms on the half-core proxy (both SMT siblings busy). Under 10 ms. Status
RUNNING: the 10^5-program worst-case search and the O-1.14 laptop relay run are owed before the row reads PASS.
What a failure moves: an acceptance-rule bound on verifier cost; the ladder's ceiling set from the measured core.
### F7. Era-draw bias harness and census (node lane harness, hash lane census)
Method: the fast-time 3-node network (`infra/fast-time/`) with an adversary withholding or publishing the last blue
block before C_era(n) to re-roll the draw; a census of 2^20 era seeds for stride, ROT and weight-perturbation
classes with gain over 1.1x; the 64-bit seeding of the day-key stream against the spec's intent. Known-failed shape:
a re-roll of the era draw inside the 2 s publish window, or an era class (stride bijection, all-equal ROT, low-weight
M) with a chip gain. Gate: no re-roll inside the publish window; no era class with gain over 1.1x at a fraction over
2^-20; the draw's input set as the spec states it.
Result (full record `docs/analysis/attack-pass/f7-era.md`; harness `tools/attack/f7-era/`). Census: 2^20 and 2^24 era
seeds through `generator::era_draw` over `V3_ALLOWED` (the chain's path), classified; the planted known-fail/known-pass
of the classifier fired and the sound draw raised nothing. No era class with gain over 1.1x at any fraction (the richest
is M = 1 at 1.0034x, absent in 2^24; every class over 2^-20 is 1.0000x to 1.0007x); the stride is a bijection on every
sample (0 even M), R and pos and the M bits uniform; the op-weight corners (15 to 31 of 75) are 1.0x against the GPU, 0
memory effect. The 64-bit day-key seeding is the spec's intent (spec 1.8.4); 2^16 days are all distinct, birthday 2^-33.
Harness: the 3-node fast-time network (`reroll.mjs`, ports 29800+, suffix 980) with an adversary holding the last block
before the cut; known-pass (`--vdf-ms 0`) fires at 1 of 6 cuts (seed = adversary block), known-fail (`--vdf-ms 5000`)
is silent at 0 of 6, both SOUND. The node has no era VDF yet (`seed_below` is a plain block hash, era-layout.md section
8), so the harness cannot show the real 2 s-window gate; the era draw's grinding resistance rests on the 1-hour VDF of
spec 4.4 (re-roll needs a 1,800x evaluator, spec 4.6 gives 300x; forge needs 20 days of 100% hash). Verdict: census PASS,
64-bit seeding PASS, harness INCOMPLETE with the written argument. Logs on igneum-build-1
`/srv/builds/igneum-wt-attack/attack-f7/census-2p24.log`, `census-2p20.log`, `reroll-knownpass.log`, `reroll-knownfail.log`.
What a failure moves: the draw procedure or the C_era cut rule; a redraw rule for the era stream.
Sub-row (a) CLOSED, 7 October 2026, 15:1x UK (the era-VDF lane, launched by the coordinator on this row's INCOMPLETE):
the era VDF is in the node (fork `era-vdf-node` 394a5902 on `release-0.3.20-node` c4459193, behind
`era_vdf_activation_daa`, never on any network until the founder sets it per network; repo master a4eaf766 carries the spec
text, the record `docs/analysis/era-vdf-2026-10-07.md` and the harness `tools/era-vdf/reroll.mjs`, which attacks the
era cut directly now that the node takes `pow_era_blocks` and `pow_era_lead` from the override file). Against the
REAL era cut (era 120 DAA, lead 20 on the fast-time file, three nodes on igneum-build-2) the harness fires with the
VDF off (6 of 6 cuts: the adversary's block is the cut block and its hash the seed, known the instant it is built) and
is silent with it on (0 of 6 across six cuts: the adversary's 5.4 to 6.6 s evaluation with the node's own code against
a 1 s block interval, the honest chain 3 to 10 blocks ahead when it published, three nodes agreeing on every era seed,
the record ready at every era start). SOUND both ways. The production delay: 517 s on the fastest prover measured
(chiavdf NUDUPL over GMP, 208.8K squarings/s on the same core) at T = 108,000,000 squarings, 259x the 2 s window.
Freeze sentence (the era-VDF lane's, carried to the plan's owner): "The era seed E_n is the output of a one-hour
verifiable delay (class-group Wesolowski, 1,024-bit prime discriminant, T 108,000,000, scheme byte 0 with the
hash-chain fallback as byte 1) over the blue blocks of the day ending at the era's cut block; the attack pass's F7
re-roll harness fires against the stand-in and is silent against the delay, so the era draw procedure and the C_era
cut rule are frozen with the VDF in the node, behind era_vdf_activation_daa, never until set per network." Open, not
a gate of F7: the verify is 22 ms with the group held, against the 10 ms target (once per 180 days per importing
node; O-4.6's reducer or GMP behind a feature, 0.3.22). Logs `/srv/builds/igneum-wt-era-vdf/ev-harness-out/
reroll-vdf-{on,off}-5.json` on build-2. F7: PASS on all three sub-rows.
### F8. Uniformity censuses (hash lane, on the box)
Method: the line-index distribution over 2^28 derivations; distinct lines per hash and per warp on 10^6 nonces of
three programs; the cross-hash item histogram of one epoch. Gate: the largest bucket within 6 sigma of uniform; no
hot set under 1% of items. Result: RUNNING. What a failure moves: the mask or the fold; packs re-cut.
### F9. Acceptance edges and header grinding (hash lane; one PC 2 job)
Method: the 39 edge disagreements reproduced and bounded; a search over 10^6 seeds for programs that pass rule (c)
with a hot set under 1%; the header-grinding search cost against its DRAM-locality gain measured on PC 2's RTX 5090
(one job through `tools/build-job.mjs`). Known-failed shape: a seed grind that steers a program to a hot cache set
for DRAM locality, or an edge where the closed-form stand-in disagrees with the live verifier in the attacker's
favour. Gate: zero passing programs with a hot set under 1%; the grinding gain under 1% of rate at any search cost.
Result: the `accept` path reproduces per-seed verdicts (genesis seed: 1 candidate ACCEPTED, bias max 54, 0
saturated). The header-grinding cost-versus-gain measurement needs a 5090. Status BLOCKED on the go decision: use
PC 2's 5090 through a relay run job only if PC 2 is online and mining is unaffected, else a rented pod under the
standing fleet budget. What a failure moves: the closed-form stand-in replaced by the live verdict at the edges; a
locality term in rule (c).
### F10. Ladder signal monotonicity (node lane)
Method: the fast-time harness with a weight that steps the ladder down and never up, and an 89% signal; the step
rule's monotonicity and its memoisation per seed block. Known-failed shape: a chip owner stepping the ladder down
(cheaper N) without the 90% threshold, or a step registered under 90%. Gate: no step without 90% over 7 windows in
either direction; a step down needs the same. Result: RUNNING. What a failure moves: the step rule's text in spec 01
before the ladder is frozen.
## Lane (d): the families re-run on class v5 (7 October 2026, evening; the coordinator's word on the founder's order)
Object: igneum-pow on branch `class-v5` at e4f1f275 (the frozen sub-version 3 017e7037 merged; the v5 chain draw is
the amended v4's instruction for instruction, generator 5, every item keyed by the window's state through the leaf
XOR before the first mixer; `V5_CLASS` = `mx8+sh256x27+state`), against the first v5 pack
`proto-cuda/packs-ca3-v5/v5-dn3-epoch0` (Devnet 3's genesis 4020cb43... as epoch and era seed, day 20,733, program id
e5a4ac5978462156, reproduced by the e4f1f275 build on box 2: the pairing). The four harnesses carried onto the
class-v5 tree in worktree `igneum-wt-attack-v5` (branch `attack-v5`), each with `--class v5` and, where the dataset
enters, `--state <IGSD1>` attaching the leaves through `with_leaves` as the CLI does; F8's traced derivation carries
the leaf XOR and validates bit for bit against `derive_items_leaves` and `Epoch::hash_warp` (p1 on the dn3 state at
4,096 nonces: 0 mismatches over 16,777,216 items and 64 warps; the flipped-state file mismatches: the known-fail).
Both boxes at nice 10 beside the release builds; the binaries run from copies in each run's scratch directory (AP-H2).
GitHub answered 403 (account suspended) from 17:2x UK, so this section lands on the box mirror (`build`, master and
attack-pass) by the coordinator's exception rule; nothing touches GitHub.
| Family | Class v5 run | Result | Verdict |
|---|---|---|---|
| F4 weak-day census | 2^24 chain days from 20,729 under `Shape::for_class(&V5_CLASS)`, box 1, 18:5x to 19:1x UTC | byte-identical to the class v4 census: M2 (DSP-bound) 0 of 2^24 days over 1.1x; M1 (LUT adders) 5,476 days, 3.264e-4, the same bounded tail, worst day 4,819,563 at cost 197 against the median 231; planted weak days fire (mul1all M2 unbounded, mulnaf 1.333x). The day-key draw depends on the mixer shape alone and v5 adds only the state flag, so identity is the expected and the measured result | PASS (v4's reading; AP-F4-1 stays the next-class item) |
| F8 hot-set gate (pre-freeze reading) | 64 seeds at 2^24, chain path, v5 with the dn3 state, box 2: hand-started at 18:47 UTC (11 seeds, killed on the coordinator's rule: every hand-started run off the boxes, loads 601 and 401), re-queued through `lease pool 64` at 19:23 UTC (17 more seeds), released at 19:4x UTC on the Counter ASIC lane's yield so the class v5 (c''') census, the 0.3.24 board's critical path, could take the pool | every seed read equals sub-version 3's seed for seed (0.9915x to 1.144x, p10 1.50x), as the v5 lane predicted: the leaves change the words, not the read addresses | READING, not the gate line |
| F8 hot-set gate, the frozen tip (THE GATE LINE) | igneum-pow class-v5 1c420786 (the 0.995 per-site floor (c''') on sub-version 3's rules; binary sha256 0f5c98dc41a1b3aa..., run from a copy); pairing: the library draws the dn3 epoch-0 program as e5a4ac5978462156, the harness validates bit for bit against the library on the dn3 state (0 mismatches on 66 validation lines); 64 seeds p2 to p65 at 2^24 nonces, chain path, the v5 dataset from v5-dn3-epoch0's state.igsd1, window-model control, build-2 under `lease pool` class v5 as two halves of 32 (ended 21:58:43Z and 22:03:21Z) | 61 of 64 under 1.2x of the window model (0.9919x to 1.144x, p75 1.0024x); 3 over, all inside the named four-seed residue and none new: p10 1.5047x (0x4018f5, 346 reads of 2^31, no predicted source), p8 1.3787x (0x839d33, 419), p4 1.2166x (0x400197, 363); p34 0.9997x under the (c''') floor; p23 1.0000x, p19 0.9997x, p15 0.9998x, p18 1.0001x, p56 1.0000x. Seed for seed the ratios equal sub-version 3's within 0.001 except where the floor moved a draw: the state leaves change the words, not the read addresses. Nothing to a chip | PASS (the known residue p4, p8, p10, attributed 8 October 2026 to per-site bucket concentration; the largest-bucket bound is a next class's item) |
| F9 exhaustion count, the frozen tip (A GATE LINE for the 0.3.24 move) | 10^5 chain-shaped seeds on the v5 chain path (era-composed class, `--chain`) with the dn3 state at igneum-pow class-v5 1c420786, pairing e5a4ac5978462156, build-1, ten chunks of 10,000 under `lease pool 4` (chunks 1, 3, 5 to 9 under class v5; chunks 0, 2, 4 re-leased under class release on the coordinator's order; 14,477 to 14,480 s per chunk, about 1.45 s per seed); binary copied into the run dir; interim line sent at 00:55 UTC (seeds drawn, 0 exhausted, 0 panics, max 30, F1 0 failures), which cleared the move; last chunk written 02:34:54 UTC | 100,000 of 100,000 seeds drawn, 0 exhausted, 0 panics, 0 past attempt index 31, max attempt index 30; histogram by attempt index (0 = accepted on the first draw) 0: 31,454; 1: 21,460; 2: 14,660; 3: 10,263; 4: 7,047; 5: 4,701; 6: 3,297; 7: 2,256; 8: 1,532; 9: 1,027; 10: 702; 11: 509; 12: 365; 13: 216; 14: 153; 15: 103; 16: 80; 17: 56; 18: 39; 19: 20; 20: 24; 21: 9; 22: 6; 23: 9; 24: 3; 25: 4; 26: 2; 27: 1; 29: 1; 30: 1; first-draw acceptance 0.3145, mean attempt index 2.185 (3.185 draws per seed), 4,862 seeds (4.86 percent) at index 8 or above, 255 (0.255 percent) at 16 or above; the 256-attempt cap and the deterministic last resort never reached. Meaning per tier: no epoch seed in 10^5 fails to draw a program, so the liveness halt of AP-F8-2 has no observed case on the frozen tip at this count (the bound it supports is under 3e-5 per seed at 95 percent, about one epoch in 33,000 at worst; a halt a node operator would see as a stuck epoch, a miner as a dead epoch, a holder as a paused chain), and the draw cost stays at about 3.2 candidates per epoch for every node | PASS (0 of 10^5; the record `f9-grind.md`, section (d)) |
| F1 shadow redundancy, the frozen tip (A GATE LINE for the 0.3.24 move) | 10^5 class v5 programs through the string-seed path (`generate_from_seed_bytes_program_class`, class V5, every candidate draw through the (c''') floor over 2^20) at igneum-pow class-v5 1c420786, pairing e5a4ac5978462156, build-1, `lease pool 16` class release (cores 8 to 23, re-leased 22:34:05 UTC), binary sha256 bb70bbf69a4b3223... copied into `frozen-1c420786-f1/bin`; 20,774 s of census (5 h 46 min; about 5 core-s per program, the (c''') draw cost), census.csv (sha256 4e34b669f2f5c680..., 100,000 rows) written 04:20 UTC on 8 October 2026 | 100,000 of 100,000 programs; instructions saved min 0.000, mean 0.623, max 4.688 percent (worst `attack-f1/95060` at attempt 0, 6,912 to 6,588 per iteration; `81748`, `66933`, `3006` at the same 4.688; next 4.311); chip-view ops saved mean 0.520, max 4.783; programs over 5 percent 0, over 10 percent 0; soundness: differential mismatches 0 of 100,000 (8 random states each), verifier mismatches 0 of 100,000; 0 panics; histogram of saved, 0.5 percent bins from 0: 55,241; 20,597; 11,762; 9,851; 1,484; 656; 259; 133; 13; 4; 0; 0; draw attempts per program: index 0 31,630, max 30 (the same shape as F9's). Against the v4 10^5 (max 5.078, the AP-F1-1 letter miss): the v5 tip's worst sits 0.39 points under the letter, the two top bins are empty, and the mean is unchanged (0.617 to 0.623), so the v5 leaves add no redundancy and remove the one letter miss. Meaning per tier: the shadow block of every drawn program stays within 5 percent of its naive count under the harness's rules and the honest compiler finds the same shortcuts, so no chip gets a shadow-side discount (a miner on a card pays the full block, a hypothetical ASIC gains nothing here) and the chain's verifier agrees with the harness on every program (0 mismatches), so no node disagrees with another on any drawn block. Harness gap and its fix: section 12 of the record (the running census was the old binary; the flush is in 18a9c04a for every census after it) | PASS (0 of 10^5 over the letter; AP-F1-1 FIXED-AND-PASSED on v5 at this count; the record `f1-shadow.md`, section 13) |
| F4 weak-day census, the post-freeze commit | 2^24 chain days at class-v5 8ca66afa (AP-F4-1 in the agreed form: cost A at most 205, k >= 1 or all-ROT-equal rejected, the forty redrawn), the harness on the agreed w32 convention (digits 0 to 31) and median 226; known-failed day 29,337 redrawn under the rule (cost 203 to 228), day 20,729 at 219 unchanged; build-1 `lease pool 12` class adv, 379 s, ended 22:3x UTC | 0 of 2^24 days over 1.1x on M1 (median 226; minimum cost 206 at day 27,016, 1.097x, one adder above the reject line) and 0 on M2; mean 225.79, sd 6.07 (pre-rule: 5.69e-4 over, min 203). The redraw rule removes the LUT tail by construction and the measurement agrees | PASS; AP-F4-1 FIXED-AND-PASSED against class v5 at 8ca66afa |
| F9 exhaustion count | 10^5 chain-shaped seeds on the v5 chain path with the dn3 state, box 1, ten parallel chunks (the chain draw costs about 2.2 s per candidate through (c''), so 10^6 is about fifty hours); killed before its first chunk closed, re-queued through `lease pool` | pending | pending the re-queue |
| F1 shadow redundancy | 10^5 class v5 programs through the string-seed path, box 1; the known firings fire under v5 (planted 50 of 256: 19.53 percent; the real block 0.000; the must-not-fire 1.157); the census killed before its end, re-queued through `lease pool` | pending | pending the re-queue |
## Operating hazards found by the pass
AP-H1 (box scratch cleaned by builds; found by F3, 7 October 2026, 10:0x UK). `infra/build-server/remote-run.sh`
line 71 runs `git clean -qfd -e target -e 'target-*' ...` on `/srv/builds/<worktree>` before every remote build, so
an untracked box scratch directory of one row (a venv, a log dir, a crate's `tools/attack/*/target`) is deleted by
the next build from any row. F3 protected its own directory through the box mirror's `.git/info/exclude`; the lane
then added `attack-*/`, `target-attack-*/`, `tools/attack/` and `.build-remote.log` to that file at 10:1x UK, after
which `git clean -fdn` on the mirror lists nothing (the clean has no `-x`, so the exclude file applies). The class
check is owed to the build-server lane: the clean line should spare a lane's declared scratch prefix (`-e 'attack-*'`
style, or read a per-worktree exclude list), and a CI check should fail a remote-run.sh whose clean line lacks it.
OPEN until that check lands (CLAUDE.md: a rule row closes only with its check).
AP-H2 (this lane's own, 7 October 2026, 13:3x and 14:5x UK, twice). Two census runs launched from the same crate's
`target/release` binary path on the box mirror: a rebuild of the crate at a new commit replaces the binary under a
run still in progress, and every chunk the run launches after that executes the new commit's code with the old run's
label (the 07a809a7 control's later chunks ran 8bdcbdd8; the ddacfbd3 class check's later chunks ran 017e7037). Both
runs were caught by their attempt histograms (attempts 32 and 35 under a cap of 32) and their contaminated chunks
discarded. Fix in the lane's launcher: `run-census-chain.sh` copies the binary into the run's own scratch directory
before the first chunk and runs from the copy, so a rebuild cannot reach a run in progress; a run's record names the
sha256 of the copy. Class check owed: the same rule for every lane's long run (the box's build runner could refuse to
replace a binary that a running process has open, or stamp the commit into the run's log at every chunk).
## Ledger rows
AP-F1-1 (hash lane; ruling asked). At 10^5 class v4 programs one program (`attack-f1/37341`) compresses by 5.078
percent (13 of 256 shadow instructions per pass), 0.078 points over the gate's first clause, on 1 of 100,000; every
other program is within 5 percent and none over 10. The saving is the same local shape as on every program (a
register written twice from one source with no write between), nothing crosses a pass, and clang -O3 removes the
same instructions from the honest kernel (IR counts match the harness on the worst programs), so a chip gains nothing
relative to a card: no shortcut. The gate as written counts honest-compiler simplification as compressibility. Two
ways to close: re-word gate (1) and row F1 to "compressible beyond the honest compiler's own simplification" (the
question is chip-relative compression), or a shadow-draw redundancy bound in the next
class (reject a shadow with over 12 peephole-removable instructions per pass, rejection about 1e-5; class v4 is on
the live vote). Ruling (coordinator, 7 October 2026, 12:3x UK): both. Gate (1) and row F1 re-worded to "no
compression of the shadow block beyond the honest compiler's simplification, measured against that compiler on the
same program" (sent to the cryptanalysis lane for the plan and the public report), under which the 5.078 percent
letter miss at honest-compiler parity is a PASS; and a shadow redundancy bound on the v5 generator's list beside
AP-F4-1 and AP-F8-1 (the generator refuses a shadow block whose honest-compiler simplification exceeds a stated
fraction; the v5 lane sets the fraction from F1's census), gated by F1's harness on 64 seeds of the v5 stream.
The fraction is 3.0 percent (v5 lane, class-v5 45e29cb0; 384 of 100,000 draws redrawn in its census, 3.8e-3, against
F1's histogram where the 3.0 to 5.5 percent bins hold 397 of 100,000); the plan's 1.1 sentence carries the number.
Status: F1 PASS; AP-F1-1 FIXED-AND-PASSED against v5 once the bound is in the v5 generator and F1's census passes.
AP-F5-1 (algorithm and hash lane, ours; the ladder lane is closed). The k about 0.33 chip-efficiency figure is
framed as a measured calibration ("the X9's core", `fud-ledger.md` M32 L172; "measured class", `ladder` branch
`docs/design/latency-ladder.md` section 5a). The Antminer X9 was withdrawn before launch and never benchmarked, so
k about 0.33 is a claimed datasheet bound, not a measurement. This also puts the merged ledger X34 ("RandomX has a
shipping chip") in question. Fix owed, held until the coordinator's research agent confirms the withdrawal and the
no-benchmark fact: relabel k about 0.33 as a claimed pessimistic bound from a withdrawn design in `evidence.md` row
17, `fud-ledger.md` M32 and the `ladder` branch; reverse X34 if the withdrawal is confirmed; keep 2.1x at k = 1 on
the GDDR7 measured anchor as the headline with the k range beside it. Re-gate after the re-cut. Status: FIXED on the docs rows (evidence 17, M32, ladder 5a on branch attack-ladder-5a d3cb17b6; attack-pass
rebased on master a3678789 after X36); FIXED-AND-PASSED once the site lane's X34/X36 served rows are confirmed in
one voice (no objection received) and the sweep is re-run against the re-cut sentence (the numbers are unchanged, so
the re-gate is the identity check and one `model.py --section chip` run against the new wording). Re-gate done 7 October 2026, 09:5x UK: the 5090
bench row still reads 5.7x / 4.1x / 3.2x / 2.1x / 1.5x (v3; v4 at k = 0.3 / 0.5 / 1 / 1.5), identity grep 0 hits
over 290 export files, the site lane confirmed the served text agrees. AP-F5-1: FIXED-AND-PASSED.
AP-F8-1 (hash lane; the generator fix is the Counter ASIC lane's on the v4 seam, routed 7 October 2026, 10:3x UK).
The class v4 item read map is not uniform. F8 phase D, one program, 2^26 nonces: the top 0.1 percent of items take
0.520 percent of reads against 0.115 percent for the uniform control (4.05x); the top 1 percent take 2.49 percent
(1.37x); one item (0xca5b92) takes 78,479 reads, 153x the mean; read site 15 feeds 6.37 percent of its reads into
that 0.1 percent in all 8 iterations; the excess grows with N as a real skew does. Sized: a chip caching the hot
0.1 percent in SRAM serves about 0.5 percent of reads from cache, so the shortcut is under one percent of rate today;
an auditor flags a non-uniform read map in a design that claims uniform random reads, and site 15's index derivation
is the cause to name. Fix asked: per-site index whitening or a rejected class above a bound. Re-gate: the top
0.1 percent within 1.2x of the control over 2^26 nonces on every one of 64 seeds, with F8's harness against the
Counter ASIC lane's branch. Phase E (the 64-program census) decides whether it is one program or the class.
Framing from the Counter ASIC lane (the generator's owner, 7 October 2026, 10:5x UK): class v4's item map is not
designed to be uniform per program. Layer 8 (spec 01 section 1.13.1) gives each load site k_off = below(3), so a
site reads the whole dataset, a half or a quarter under the era's stride and interleave; a quarter-window site
concentrates 4x on its quarter by design, which is the 4.05x at the top 0.1 percent, and the windows exist so a
chip's SRAM mirror must hold the whole dataset every hour (the Counter ASIC 2.0 windows-union census). The right
control is therefore the window model from the program's own 16 draws, reported beside the uniform control (what an
auditor sees first); the number that must be explained is the single item 0xca5b92 at 153x the mean (window
coincidence under the era mapping with a stated tail, or a low-entropy index source at site 15, which would be a
fault). The lane reproduces with F8's harness on branch `ca3-v4-uniform`, waits for phase E, re-prices the chip
consequence (a 0.1 percent hot-set cache, about 1.7 MB of SRAM, serving 0.5 percent of reads: under one percent of
rate) and changes the generator only on a fault beyond the model, since v4 is on the live devnet's vote. F8 was
re-briefed to carry both controls and the per-site table. Raised to the coordinator: plan 1.4 gate (4) and row F8
say "within 6 sigma of uniform"; if the design is windowed, the gate text must say "uniform within the window model
of spec 1.13.1" before the freeze tag, or every reviewer files the windows as a finding on day one.
Coordinator's ruling (7 October 2026, 11:0x UK), accepted: the right null is the window model derived from the
program's own draws; F8 is re-gated against it, and the finding stays open only for the excess beyond the window
model (the 153x item, or a low-entropy source at site 15 if the 64-seed census shows one). No generator change to
class v4 is allowed: it is on the live devnet's vote, and a class change before the flip splits the chain. If the
census shows a real fault it goes to the coordinator priced; otherwise the record carries the documented null and
the hot-set bound (a 0.1 percent cache, about 1.7 MB of SRAM, under one percent of rate) goes into the next class.
Gate wording settled (coordinator, 11:2x UK): plan 1.4 gate (4) and row F8 now read "uniform within the window
model of spec 1.13.1, layer 8; the excess beyond it within 6 sigma over 64 seeds", carried into the plan's scope
text by the cryptanalysis lane so the public report states the windows.
Mechanism (hash lane, branch `ca3-v4-uniform` 095f84a7, `docs/analysis/ca3-v4-uniform.md`, harness
`tools/ca3-v4-uniform`, 7 October 2026, 12:3x UK): the windows-union null (a Poisson mixture at 416 / 288 / 736 / 608
reads per item by quarter from the program's 16 draws) moves the top 0.1 percent from 0.115 to 0.160 percent, 1.39x,
not 4.05x; every per-site row of F8's attribution except site 15 is the window model. The rest is the LOAD SOURCE:
site 15 is the load at 63 reading r6, whose last writer is `or` at 61 (r6 = r6 | r4), so the source is all-ones with
probability about (3/4)^32 per read; under the era map x = 0xffffffff is item 0xca5b92, the hottest item exactly, and
the next seven hottest are the seven one-zero-bit sources whose zero survives the window mask (7 of 7); the measured
count fixes the bias at p = 0.7585 per bit. The class: a load whose source's last writer is lossy (or: 0.30 percent
of a site's reads on 0.1 percent of values; mul, trailing zeros: 1.07; mulhi: 0.79; an or of an or: about 4.5).
Static census of 1,024 chain-shaped v4 programs: 96.6 percent carry a lossy-sourced load (48.5 percent or, 4.9
percent an or chain, 73 percent mul, 64 percent mulhi); predicted S_0.1 median 0.45, 90th 0.88, 99th 5.3, max 9.8
percent; p1 / p2 / p3 predicted 0.58 / 0.32 / 4.72 against measured 0.52 / 0.27 / 4.60. The fault sits in the
acceptance rule's blind spot: part (a) takes any write as fresh, part (c) counts saturation on final values only.
Consequence: the 1.2x-against-window gate fails 96.6 percent of today's programs, so it is withdrawn as a v4 gate and
becomes the v5 generator item's gate (draw a load's source from registers whose last writer injects; a dynamic check
counting saturated load sources), with F8's phase E as its test. Chip side: the top 0.1 percent of items is 1.07 MB
of SRAM (0.53 mm^2, about USD 0.25) serving 0.52 percent of p1's reads and 4.6 percent of p3's, at most 1.005x and
1.048x in rate; the ceiling under rule (c)'s 120-of-128 floor is one site repeating its item in all 8 iterations,
6.25 percent of reads, 1.067x. That 1.067x is the v4 hot-set bound the record carries. No generator change to v4;
the hash lane takes the two flip options priced to main.
Ruling (the founder, 7 October 2026, 15:2x UK): option A, the class v4 amendment ships in 0.3.20, the feature node (0.3.19 is the app-only cut on the unchanged 0.3.17 node pin; corrected by the coordinator) (a load's source drawn
only from registers whose last writer injects or is a rotate, the v5 rule applied now; a new program stream and
seven re-exported packs on branch `ca3-v4-amend`, the hash lane), with limited testing. This lane's part is the proof
of the fix: F8's hot-set census at 2^24 nonces on each of 64 seeds of the amended stream, on the box's CPU path as
phase D ran, gate: the top 0.1 percent of items within 1.2x of the window model derived from each program's own 16
window draws, one number per seed; reported to the hash lane, main and the Counter ASIC lane. The amended stream has
no lossy-sourced load by construction, so a seed over 1.2x there is a finding against the model's own tail, not the
fault, and the record says which.
Second harness (F9 sub-row b, 10^6 seeds, 12:5x UK): the same class from the other side, the per-site address trace:
11,696 of 1,000,000 passing programs concentrate 1 percent or more of their reads on a hot set (worst 17.3 percent,
seed 842871, an `or`-written load source all-ones in 36 percent of evaluations), so F9-1 merges into AP-F8-1 and
F9's harness is the second re-gate of the amendment, run on the amended stream beside F8's 64-seed census.
Re-gate interim (7 October 2026, 13:2x to 13:5x UK, box 2): F8's 64-seed census at 2^24 nonces against the amended
stream (igneum-pow 8c728ca3, sub-version 1; pairing verified, the harness draws the devnet epoch-0 program as
1a4230699a6b9c60) at 30 of 64 seeds shows nine over 1.2x of the window model (p31 29.27x, p11 5.45x, p19 3.32x, p6
3.11x, p23 2.04x, p4 1.57x, p10 1.50x, p26 1.30x, p25 1.28x), p6's hottest item predicted from "site 13, r0,
all-ones, last writer a load at 12": a load-after-load chain (a hot address yields a fixed dataset word, which is the
next load's address), which the source rule admits because a load injects. Rule-level reading, checkable in code:
generator.rs line 1326 sets `entropy_kept[dst]` true for a rotate whatever it rotated, so an or-saturated register
rotated once is an admitted source and the rotate preserves the saturation. RETRACTION: F9's hot-set census run on
box 2 against the 8c728ca3 build (10^6 seeds, 1,871 flagged, worst 9.66 percent) was not a re-gate: the F9 harness
draws through `candidate_class` with its own era class, not through `chain_program` where the rule lives, and the
amended and the old binary print the identical program for seed igneum-f9/518927; those numbers describe the old
stream under a changed evaluation and are withdrawn; the harness is being given a `chain_program` draw mode so it can
serve as the second re-gate. The hash lane confirmed the reading (14:0x UK): the amendment's rule is keyed on the
era-composed class, so a draw with no era (F9's path) is the old stream, and on the chain path the residual is real:
p6's load at 12 had a saturated source itself, read one constant word and left a constant in r0, which the rule
counts as injecting; a rotate keeps 0xffffffff, so or-then-rotate-then-load passes too. Both are saturation delivered
through a writer that preserves it. Fix shape put to the owner of sub-version 2 (the Counter ASIC lane): dataflow
freshness instead of a one-writer look-back (fresh at the start; a load keeps dst fresh only if its source was fresh;
add, sub, xor, mad, shfl fresh if either operand was; rotl, rotr only if the operand was; or, mul, mulhi never; a
load's source drawn only from fresh registers), with the dynamic (c') check on load sources as the backstop; a stream
change, so sub-version 2 with new packs, ids and fingerprints.
RE-GATE VERDICT on sub-version 1 (7 October 2026, census ended 12:55:55 UTC, 13:55 UK; box 2; igneum-pow 8c728ca3
paired with release-0.3.20-node 8097d600, pairing id 1a4230699a6b9c60 verified; 64 seeds p2 to p65 at 2^24 nonces,
chain path, window-model control; log `/srv/builds/igneum-wt-attack-regate/attack-f8-regate/log/`): FAIL the pass
line. 53 of 64 seeds under 1.2x of the window model (0.9915x to 1.16x, no predicted source); 11 over:
| Seed | Over the window model | Over flat | Hottest item, reads of 2^31 | Predicted source |
|---|---|---|---|---|
| p31 | 29.27x | 31.99x | 0x74e2b8, 5,365,527 | site 4, r4, all-ones, last writer rotl at 3 |
| p11 | 5.45x | 6.00x | 0x0eec66, 31,486 | site 1, r7, all-ones, last writer or at 63 (the previous iteration) |
| p45 | 4.55x | 6.37x | 0x400000, 5,644 | site 1, r4, zero, last writer mulhi at 59 |
| p19 | 3.32x | 3.93x | 0x400000, 28,114 | site 37, r5, zero, last writer load at 32 |
| p6 | 3.11x | | 0x3bf40d, 13,792 | site 13, r0, all-ones, last writer load at 12 |
| p23 | 2.04x | 2.85x | 0x09dd36, 13,848 | site 16, r7, all-ones, last writer load at 14 |
| p4 | 1.57x | 2.11x | 353 reads | none (window tail) |
| p34 | 1.51x | 1.87x | 0x400000, 1,547 | site 23, r6, zero, last writer rotr at 12 |
| p10 | 1.50x | 1.65x | 353 reads | none (window tail) |
| p26 | 1.30x | 1.54x | 0x000000, 7,637 | site 10, r1, zero, last writer rotl at 2 |
| p25 | 1.28x | 1.65x | 363 reads | none (window tail) |
Three residual classes, each a constant (all-ones or zero) delivered to a load through a writer the rule admits:
(1) saturation or zero preserved through rotl, rotr, load or mad; (2) zero made by mulhi; (3) the iteration
boundary, where the rule's writer state starts fresh at instruction 0 so an or at 63 feeds a load at 1. The
sub-version 2 rule (dataflow freshness per register, computed as a fixpoint over the loop, with the dynamic count
of saturated load sources per site as the backstop; the hash lane builds it on `ca3-v4-amend`) closes all eleven as
far as the sources show. Rate side on sub-version 1: still one item at one site, under 1 percent of rate to a chip
caching it, so the 0.3.20 ship is safe on rate; the auditor's flag is what sub-version 2 removes. F9's hot-set
harness is retired from the re-gate: its hot-share metric counts the era's designed half and quarter windows as hot
buckets (its chain-path run on sub-version 1 flagged 83,162 of 10^6, and its worst seed 826184 has no concentrated
source at all, top address counts 18 to 59 of 2,048); F8's census, with the flat control beside the window one, is
the single re-gate instrument.
Sub-version 2 (07a809a7, the stream identical at 8bdcbdd8 for every seed accepting within 32), the same 64 seeds at
2^24, 13:10 to 14:2x UTC: at 39 of 64 seeds, 8 over 1.2x of the window model, worst p23 4.82x. Three are the window
model's tail (p4 1.22x, p8 1.38x, p10 1.50x, no predicted source); five are constants the freshness rule cannot see
because it tracks lineage, not value: p23 (0x000000, 41,727 reads, zero from xor of a register with itself at
instruction 0), p34 1.25x (sub of a register with itself), p15 2.57x (zero through rotl at 0), p18 2.50x and p19
3.32x (a load whose address is constant delivers one word to the next load; p19 is byte for byte the sub-version 1
program). (c') cannot catch them: 164 of 16,384 per site is about fifty times coarser than the gate (p23's item is
0.002 percent of all reads and still 4.8x at the top 0.1 percent). Fix shape sent to the hash lane: forbid
self-operands for xor, sub and mad in the draw; a dynamic per-site bound on the most repeated source value (any
value) set from the gate; a load's dst fresh only if its source passes it. Rate side unchanged (one item at one
site, nothing to a chip); the auditor's uniformity test is what fails.
Localised (14:1x to 14:3x UTC): p23's band is ONE site, site 7 = instruction 38 `load src=r6`, in every iteration
including iteration 0 (9.5 percent of that position's reads on the top 0.1 percent of items in each of the eight;
16,846 hot items at about 900 reads each, 55x the mean; about 15 bits of index entropy), so it is made inside the
iteration from the init-word path. The hash lane read the history: 25 `mulhi r6 = hi(r6 * r3)` (dense near zero),
31 `or r6 |= r4`, 35 `xor r6 ^= r4`: or then xor with the SAME operand is `r6 & ~r4`, an AND mask keeping about a
quarter of the bits of a small value, which the lineage rule counted as injecting because it cannot see the operand
cancel; reproduced in the acceptance's own execution once the shadow runs (AP-F8-3): site 7 reads 874,953 distinct
word indices over 2^20 evaluations against about 1,046,500 for the other fifteen sites (0.84 of uniform, 2.2 s) and
0.55 at 2^24 (35 s). Neither dataset- nor nonce-dependent: a rule reaches it. Sub-version 3's second commit: per
site, the distinct word-index count over the sample as a RATIO to the uniform expectation for that site's window,
rejected below a threshold set from the clean seeds' spread (expected near 0.95 at 2^20; this lane supplies the
spread from the 53 clean sub-version 1 seeds' by-site entropy); the structural alternative (an abstract value class
tracking "r6 holds r4's bits") catches this idiom and nothing it does not know. Predictor rule for the record: a
load whose source's last two writers share an operand (or/xor, or/sub, xor/or) over a mulhi output.
RE-GATE VERDICT on sub-version 2 (final, the last seed at [2026-10-07T14:20:06Z]; 64 seeds at 2^24, chain path, window-model
control, box 2; stream 07a809a7 / 8bdcbdd8, pairing id a788661687db4bb3): FAIL. 55 of 64 under 1.2x (0.9915x to
1.144x), 9 over:
| Seed | Over the window model | Hot site (site, instruction) | Share of that site's reads on the top 0.1 percent | Bucket entropy of uniform | Predicted source |
|---|---|---|---|---|---|
| p23 | 4.82x | 7, 38 | 9.43 percent | 0.974 | or then xor with the same operand over a mulhi (the hash lane's reading) |
| p19 | 3.32x | 15, 62 | 6.64 percent | 0.964 | zero through a load (unchanged from sub-version 1) |
| p15 | 2.57x | 2, 12 | 4.49 percent | 0.982 | zero through rotl at 0 |
| p18 | 2.50x | 6, 30 | 5.55 percent | 0.937 | all-ones through a load |
| p56 | 2.01x | 2, 10 | 3.34 percent | 0.994 | unattributed (new over sub-version 1) |
| p10 | 1.50x | 8, 28 | 2.04 percent | 0.979 | bucket concentration at a narrow-window site (r0, window 2^22, offset 1, last writer mad at 20): largest 256-item bucket 5.6x window expectation, index entropy 13.71 of 14 bits, hottest 0x4004da at 362 reads, source none (identical to sub-version 1) |
| p8 | 1.38x | 14, 51 | 1.42 percent | 0.980 | bucket concentration at a narrow-window site (r7, window 2^22, offset 2, last writer xor at 44): largest bucket 3.1x, entropy 13.72 of 14; plus site 6 (instr 33, r3, window 2^23, mad at 30) at 0.834 percent, bucket 3.5x; hottest 0x837de4 at 420 reads, source none |
| p34 | 1.25x | 1, 13 | 1.35 percent | 0.997 | one-bit value through sub (r3, window 2^23, offset 1, last writer sub at 5): largest bucket 3.5x, entropy 14.96 of 15, hottest 0x800010 at 541 reads, saturated source 0.0001 percent |
| p4 | 1.22x | 1, 8 | 1.45 percent | 0.981 | bucket concentration at a narrow-window site (r2, window 2^22, offset 1, last writer mad at 4): largest bucket 4.5x, entropy 13.74 of 14, hottest 0x4000e7 at 355 reads, source none (1.57x on sub-version 1) |
Every failing seed is one low-entropy load site. Clean-seed spread of the per-site bucket entropy (848 site rows of
sub-version 1's 53 clean seeds): min 0.9865, p1 0.9961, p5 0.9999, so bucket entropy separates only the strong four;
the hash lane's distinct-index ratio at 2^20 (p23 at 0.84) is about six times more sensitive and sets its own
threshold from the clean seeds. Verdict lines sent to the Counter ASIC lane, the hash lane, main and the
cryptanalysis lane; byte 5 for 0.3.21 stands on this evidence.
Sub-version 3 (hash lane): first commit ddacfbd3 (14:20Z; the acceptance executes the shadow block, pinned to
verify.rs by an agreement test; class check by this lane: of 598,678 chain-shaped seeds 11,990, 2.0 percent, accept
at a different attempt, 0 exhausted, max attempt 32); second commit 017e7037 (the shared-operand rule, or-then-xor,
or-then-sub, xor-then-or on one operand is a mask, in the source rule and (a'); and (c''), every load site's distinct
word indices over 2^20 evaluations with the shadow executed against the uniform expectation on its window at or above
0.98, the last test of the chosen candidate). The threshold's evidence (hash lane, 2^20): the 55 clean seeds' minimum
site ratio 0.9960, p1 0.9990, median 1.0000; the strong five p23 0.8361, p18 0.9274, p19 0.9335, p15 0.9432, p56
0.9654; floor 0.98 sits 0.015 from each side. At 2^24 the weak four (p34 0.9181, p4 0.9614, p8 0.9630, p10 0.9612)
share their value with two clean seeds (p44 0.9612, p52 0.9613), so the 2^24 stage is not taken and p4, p8, p10 and
p34 are the tail. The tail attributed (hash lane, 8 October 2026, 09:40 to 09:46 UK, attack-f8 census at 2^24 with
the window-model control and by-site attribution, tree b38b4af6 on the frozen 017e7037; the gate ratios reproduced to
four places, p4 1.2169x, p8 1.3774x, p10 1.5036x, p34 1.2501x, the hot-set verdict clear on the windowed control for
all four): each is a per-site bucket concentration of about a quarter bit (0.26 to 0.29 bits short of 14 on a 2^22
window; p34 0.04 of 15) at one narrow-window load site whose last writer is mad, xor or sub, with every other site at
its flat share; the ratio tracks the largest-256-item-bucket excess (5.6x gives 1.50x, 3.1x to 4.5x give 1.22x to
1.38x); (c'') passes them at 0.9927 to 0.9963 because distinctness does not see a bucket; the check that would catch
all four is a per-site largest-bucket bound (about 2x window expectation at the 2^20 units), a generator change for a
next class, never for the frozen ones. Meaning per tier: a quarter bit at one site is under the window model's own
spread (the gate line's 61 of 64 stands), so no card or chip gains a cacheable hot set from it; the bound is the next
class's item, not a change to 017e7037 or 1c420786. The ratio refuses about 4 percent of candidates that pass every other test
(4,099-program census at 017e7037: mean attempts 2.086 against 1.998, max 17, 0 lossy-sourced load sites of 65,584,
0 exhaustions; suite 103 of 103; devnet epoch-0 at attempt 1, id a785001687d8688a, pairing verified by this lane).
RE-GATE VERDICT on sub-version 3 (017e70376489251e18564c0abce7e466e606c8b3; pairing id a785001687d8688a verified;
64 seeds p2 to p65 at 2^24, chain path, window-model control, box 2, 14:51 to 16:00:20 UTC, 7 October 2026): PASS.
60 of 64 under 1.2x (0.9915x to 1.144x); the four over are the named tail, attributed 8 October 2026 (per-site bucket concentration, the AP-F8-1 tail paragraph): p10
1.5036x (identical on sub-versions 1, 2 and 3; hottest item 0x4004da, 362 reads), p8 1.3776x (0x837de4, 420), p34
1.2505x (0x800010, 541, the one-bit value through sub at 5), p4 1.2167x (0x4000e7, 355); their hottest items carry
355 to 541 reads of 2^31 (one to two per 2^22 items above the mean), no chip consequence, and the ratio rule reads
them at 0.9927 to 0.9963 at 2^20, inside the clean spread. Every strong seed of sub-versions 1 and 2 is under the
line (p23 4.82x to under 1.2x, p19, p15, p18, p56 likewise). Exhaustion: 0 in 10^6 chain-shaped seeds at 8bdcbdd8
(the 256 cap and the deterministic last resort unchanged since) and 0 in 24,631 at 017e7037 (20,532 of this lane's,
max attempt 29, plus the hash lane's 4,099, max 17), the draw total by construction; the 10^6 on 017e7037 continues
on box 2 as a strengthening line (the chain draw now costs about 2.2 s per candidate through (c''), so about two
days) and is not a condition. Node consequence, not a gate: about 2 attempts at 2.2 s each per epoch per node, 4 to
5 s at one epoch an hour. Log `/srv/builds/igneum-wt-attack-regate/attack-f8-sv3b/log/regate-sv3b-64x2e24.log`.
Status: FIXED-AND-PASSED. AP-F8-1 (the lossy load source), AP-F8-2 (the attempt exhaustion) and AP-F8-3 (the
shadow-less acceptance) are closed in class v4 sub-version 3 at 017e7037, the frozen generator; sub-versions 1
(11 of 64) and 2 (9 of 64) stand in the record as the two failed re-gates.
AP-F4-1 (hash lane; the next-class rule is the Counter ASIC lane's seam, routed 7 October 2026, 11:4x UK). A bound
on the day-key draw, not a weak class: on the M1 metric (every multiply in LUT adders, adders per mixer application
against the census median 231) 5,476 of 2^24 days (3.26e-4) and 87,426 of 2^28 (3.26e-4) gain over 1.1x, the tail
of a sum the exact convolution predicts to 0.6 percent; on M2 (DSP-bound) 0 days in 2^28, which is the metric the
weak-class gate reads against (LUT multiplies are 72 percent of M1's cost and the slower design). Worst day in 2^24:
chain day 4,819,563 (NAF sum 149, cost 197, 1.173x); worst in the public calendar: chain day 29,337 (23.6 years in,
NAF sum 158, cost 206, 1.121x, M2 1.000x), reproduced through `igneum-pow export` (memhard.h equal to the harness).
Priced: at most 12.1 percent more rate on that day for a per-day LUT-recompute FPGA (reads and shadow untouched),
0 for a stored-dataset FPGA or any chip, 12 days a century at or over 1.1x (0.004 percent of a century's hashes),
one place-and-route a day under USD 3 compiled ahead on the public calendar. Remedy for the next class, class v4
untouched: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values,
plus NAF weight at least 4 per word and at least 4 distinct ROT amounts; rejection 6.1e-4 per day; first calendar
redraw day 22,633; no pack changes. Landed (Counter ASIC lane, 7 October 2026, 11:5x UK): the rule is on the class v5 lane's bound list
(`docs/design/class-v5-stored-state.md` section 11) with F4's harness as its gate, re-gated by this lane against the
v5 branch once its `accept.rs` carries it. The brief's rank 3 (funding.md B2, the untested all-equal ROT draw of
MEMHARD.md) now reads "a bounded tail, measured", with the F4 record as the source.
Reconciled with adv-mixer-2's independent census (7 October 2026, 21:5x UK; F4 record section 9): the same class
and cost form; this record's NAF counted the carry digit at position 32, which a 32-bit multiplier never pays, so the
agreed figures are adv-mixer-2's: median 226, a 1.1x gain at cost A at most 205, 5.69e-4 of days (2^-10.8), 15 days
a century, worst 2050-04-28 (day 29,337) at 1.113x; the DSP-bound readings agree at 0; the redraw rule for the next
class takes adv-mixer-2's form (cost A at most 205, or k >= 1, or the eight ROT equal).
Status: F4 PASS against v4; AP-F4-1 FIXED-AND-PASSED against class v5 at 8ca66afa (7 October 2026, 22:3x UTC: 0 of
2^24 days over 1.1x on both metrics with the agreed rule in the draw).
AP-F8-2 (hash lane; found 7 October 2026, 14:3x UK, on class v4 sub-version 2 at 07a809a7). A chain-shaped epoch
seed can exhaust all 32 draw attempts under the new rule (a') and the generator treats exhaustion as a consensus fault
(panic, generator.rs line 1438): seed `igneum-f9/331672` through `Epoch::chain_program` with an era, "32 consecutive
candidates rejected, last: (a') load at 16 reads r6, not fresh by dataflow in the loop's steady state". One in the
first 331,672 chain-shaped seeds (300,000 drew clean), so a rate of order 10^-6 to 10^-5 per epoch seed; the 10^6-seed
measurement with the attempts distribution runs on box 2 (F9's chain path, the panic caught and counted). Meaning:
an exhausted epoch seed is an epoch no node can draw a program for, a liveness halt, and the seeds are VDF outputs
nobody can steer around it; at one epoch an hour the bracketed rate is one halt per 11 to 40 years, which the public report's readers
would compute from the rule as written. Sub-version 1: 0 exhausted in 10^6 chain-shaped seeds. Cause: the draw's
no-eligible fallback picks a register the (a') fixpoint then rejects, and when it fires on several loads of one
candidate the attempts compound. Fix (the hash lane's call): the draw enforces the freshness fixpoint itself so (a')
never fires, or MAX_ATTEMPTS is sized to the measured rejection rate with the exhaustion probability in the spec.
Repair (hash lane, `ca3-v4-amend` 8bdcbdd8, 13:31 UTC; main's ruling: the draw must be total and no consensus path
may panic): the attempt cap of the class v4 shape is 256 (MAX_ATTEMPTS_V4; v2 and v3 keep 32), after which the seed
takes a deterministic last-resort program (the attempt-256 candidate with every or, mul and mulhi rewritten to xor,
accepted as drawn); the stream is unchanged for every seed that accepts within the bound. Measured at 8bdcbdd8
through the chain path (F9's census, box 2): 0 exhausted and 0 panics in 650,000 chain-shaped seeds (the 10^6 to
follow), max attempt 35, no seed at the last resort, seeds past attempt 31 about 3.2e-6 (5 in 1.55 million draws,
inside the (2/3)^32 = 2.3e-6 estimate), per-attempt rejection 0.67 (attempt histogram 232,235 / 155,322 / 103,509 /
68,858 / ...), mean about 2 attempts per seed; seed 331672 accepts at attempt 32. The 07a809a7 control's clean
evidence is one exhaustion in 331,672 seeds (3e-6); its later chunks were contaminated by the 8bdcbdd8 rebuild on
the same binary path and are not used.
Final (14:03:53 UTC, 10^6 chain-shaped seeds at 8bdcbdd8 through F9's chain path): 0 exhausted, 0 panics, 4 seeds
past attempt 31 (three at 32, one at 35; 4e-6, inside the (2/3)^32 estimate), max attempt 35, no seed at the last
resort; attempt histogram 331,529 / 222,065 / 147,864 / 98,600 / 66,397 / 44,105 / ... / 1 at 31 / 3 at 32 / 1 at 35,
a per-attempt rejection of 0.67 and a mean of 2.0 attempts per seed. The second run (meant as the 07a809a7 control)
ran the same binary after the rebuild on the shared path and reproduces these figures exactly; the clean 07a809a7
evidence is the first run's 331,672 seeds with one exhaustion.
Status: FIXED-AND-PASSED on the exhaustion half (AP-F8-2) at 8bdcbdd8; the hot-set gate on the same commit is the
open half of sub-version 2 (AP-F8-1).
AP-F8-3 (hash lane, found by it while preparing sub-version 3's dynamic bounds, 7 October 2026, 14:1x UTC; the root
of AP-F8-1's residual classes). `accept.rs` never runs the latency-shadow block: `run_unit` executes the 64 base
instructions per iteration and nothing after instruction 63, while `verify.rs` and every kernel run the shadow 27
times at the end of each iteration. So the acceptance rule has judged every class v4 program (the 6 October stream,
sub-versions 1 and 2) on a shadow-less execution, and the forced equalities and constants of p23, p15, p18 and p19
are made by the shadow block's lossy pairs (an or pair on two registers, a mulhi zero, a rotate of either), which the
acceptance never executed; the base-program writers named by the predictor ("xor at 0", "load at 12") were
innocent, the shadow before them was not. Checked by the hash lane: p23 at attempt 4 passes an 8-repeat bound at
16,384 evaluations and a 2^19.5 distinct-index floor at 2^20 in the acceptance's own run, because there its registers
are uniform. Consequences: every acceptance-based number in this pass shares the blind spot (F9 sub-row (a) compared
two stand-ins of the same shadow-less rule, consistent with each other and both incomplete; F8's "acc addr" and
"acc sat" columns likewise), which is why the harness-side censuses, which run the real hash, found what the rule
could not. Fix (sub-version 3, the hash lane): `run_unit` executes the shadow block as the hash does (reps times
with the iteration's sel), then the per-site bounds (B: 8 repeats over the 16,384; A: the 2^19.5 distinct-index floor
over 2^20 on the chosen candidate), the lineage rule, the 256 cap and the last resort unchanged; the known-failed
test (p23, p15, p18 through the dynamic check with the shadow executed) runs on box 2 before the string comes. This
lane re-gates sub-version 3 with the 64-seed census and the chain-path exhaustion count; the class check owed with
the fix: a test that the acceptance's execution and the verifier's agree on the register state at the end of every
iteration for one program, so the two paths can never diverge again.
Status: FINDING-OPEN; closes with sub-version 3's re-gate.
Any further finding is logged here and in `docs/fud-ledger.md` with its owning lane (hash and algorithm: fixed in
`igneum-pow` behind a test and re-gated; node: the node lane, relay agent) before the row is marked FIXED-AND-PASSED.

View file

@ -0,0 +1,344 @@
# Attack pass F1: shadow block compressibility and shortcut search
Row F1 of `docs/plans/cryptanalysis.md` section 4.2, fed into `docs/analysis/attack-pass-2026-10.md`.
Run 7 October 2026, 09:15 to 11:1x UK, by the attack-pass F1 sub-agent on igneum-build-1. Times to humans UK;
log lines UTC. Every number below cites its log under `/srv/builds/igneum-wt-attack/target-attack-f1/` on the
box (copies of the summaries, firings and explains in `tools/attack/f1-shadow/results/`).
## 0. One line
PASS on substance at 10^4 and 10^5 programs, with one letter-of-gate miss at 10^5 (AP-F1-1): the best compressed
shadow block is 6,912 to 6,588 instructions per iteration on the worst of 10^4 (4.69 percent, seed
`attack-f1/8556`) and 6,912 to 6,561 on the worst of 10^5 (5.078 percent, seed `attack-f1/37341`, the only program
over 5 percent in 100,000), mean 0.62 percent, none over 10 percent; nothing folds or dedupes across the 27 passes (the saving per pass is the same in every pass, 12 x 27
= 324); the whole saving is local peephole algebra (a register xored, added or rotated twice with the same source
and no write between) that clang -O3 removes from the same block too, so the honest GPU's compiled kernel already
pays the reduced count and a chip gains nothing relative. Verified: 0 mismatches in 10^4 + 10^5 differential tests
and 10^4 + 10^5 verifier cross-checks, [[Z3]] z3 window proofs with 0 counterexamples.
## 1. Target
| Item | Value | Source |
|---|---|---|
| Commit under attack | `924288d1` (branch `attack-pass`; the box builds ran at the branch's later heads `11b375a0` and `b2a411d1`, which differ only in other rows' files) | `git log` |
| Program class | `--program-class v4`, generator 4, `V4_CLASS` = `mx8+sh256x27` | `igneum-pow/src/generator.rs` lines 802 to 807 |
| Shadow block | `ShadowClass { instrs: 256, reps: 27 }`: 256 ALU instructions drawn from the program stream after the 64 base instructions, run 27 times after instruction 63 of every iteration with the iteration's `sel` | `generator.rs` lines 355 to 376 and 1257 to 1290; `verify.rs` lines 383 to 388 |
| Shadow instructions per hash | 8 x 256 x 27 = 55,296 | `ShadowClass::instrs_per_hash` |
| Shadow op families and weights (of 75) | add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 | `NONLOAD_WEIGHTS`, `generator.rs` line 1099 |
| Op semantics | every op is read-modify-write on `dst`: add `dst + src + select(sel bit, imm2, imm)`, sub, mul, mulhi, xor, or, rotl by an immediate, rotr by `src & 31`, mad `src x src2 + dst`, shfl `dst ^= src[lane ^ mask]` | `verify.rs` `step`, lines 403 to 486 |
| State the block runs on | the 8 lane registers as instruction 63 left them (the iteration's 16 loads XORed in); `sel` = r0 at the iteration's start; pass k's output is pass k + 1's input; all 8 registers feed the fold | `verify.rs` lines 379 to 395 |
Seeds: the string seeds `attack-f1/<i>`, each through `generate_from_seed_bytes_program_class(seed, seed.as_bytes(),
ProgramClass::V4, None)` (the acceptance rule's redraw included). Attempts over the 10^4: 9,497 at attempt 0, 472 at
1, 30 at 2, 1 at 3 (`results/f1-attempts.txt`), the 5.0 percent rejection rate of spec 1.4.6.
## 2. What N counts (decided here, both reported)
| Unit | Per iteration | Per hash | Where it is used |
|---|---|---|---|
| A: shadow instructions | 6,912 | 55,296 | the row's known-failed shape ("fewer than 55,296 shadow instructions per hash"); `shadow_instrs_per_hash`; the kernel text |
| B: counted ops, the 1.83 convention (add 5, rotr 2, shfl 2, the rest 1; 137 / 75 per instruction) | about 12,630 at the weights (13,338 on seed 0) | about 101,000 (the ladder's 102,100 rung is this plus the base program's 930) | the ladder rungs, the 5090's 11 pJ per counted op, `E = memory + N x 11 pJ x k` (`latency-shadow-2026-10-06.md` section 6, `algorithm.md` 5.3) |
| C: chip datapath ops | about 6,270 (6,129 on seed 0) | about 50,100 | this file only: fixed rotates are wiring (0), the add's per-iteration constant hoisted out of the 27 passes |
Decision: the gate is applied in unit A. (1) The row's own failed shape is written in instructions. (2) Unit B's
extra 0.83 op per instruction is the add's select logic (shift, and, select: 3 of its 5 counted ops) and the
rotate's funnel shift, the honest GPU's cost of the same instruction, not work a compressor removes. (3) The chip
model's `k` floor is derived per instruction (`algorithm.md` 5.3: 0.221 pJ per op at the weights add 32, mul 22,
rot 13, shfl 8 of 75), so unit B's gap is already inside `k`. Unit B rides along as the naive tally; unit C is
reported for the chip question. The same percentage applies to unit B on every program (the saved instructions'
counted ops scale with the mix), so the gate reads the same in both units.
Unit note for the algorithm lane (AP-F1-1, below): the `k = 0.3` floor divides a per-instruction energy by a
per-counted-op energy.
## 3. Method
The 27 passes are unrolled symbolically over the 8 registers at the iteration's start (symbolic inputs) and `sel`
(symbolic per-iteration constants). Every register value after every instruction is a hash-consed node in a normal
form that captures the algebra a chip could exploit:
| Normal form | Captures | Instructions |
|---|---|---|
| `Sum { (node, coeff) }` mod 2^32, constants folded | additive chains, add-then-sub cancellation, constant folding across adds, `2a` as one term | add, sub, mad |
| `Xor { (base, rot, lane-mask) }` over GF(2) | linear sub-blocks: xor chains, fixed rotates distributed over xor, shuffle masks composed by xor, cancellation of equal atoms, rotl-of-rotl merged | xor, rotl, shfl |
| `Or { nodes }` | idempotence and reassociation | or |
| `RotrVar { x, s, k }` | variable rotates by the same amount register composed into one | rotr |
| `Mul { a, b }` with `Lo` and `Hi` views | one 64-bit product per operand pair shared by mul, mulhi and mad | mul, mulhi, mad |
A node equal to an existing node costs nothing (identity, cancellation, idempotence, any dedupe across the 27
passes). Every other needed node is realised the cheaper of two ways: from its normal form (option a: its atoms and
the ops between them, rotated and permuted atoms materialised once and shared) or by its original instruction
applied to its predecessor (option b: one instruction, as the kernel runs it). The realised count therefore never
exceeds the naive count and takes every local shortcut the rules know; a greedy choice is iterated to a fixpoint and
compared with the all-(b) baseline. Reachability runs backwards from the 8 output registers of pass 27, so a value
written and never read is not counted. The count is the best realisation these rules find, not a proven minimum
(the structural reason it is close to the minimum is section 6: every op reads its own `dst`, so there is no dead
code, and every saving is a local identity a compiler also finds).
Soundness, three ways: (1) every program's normal-form DAG is evaluated concretely on random 32-lane states and
compared with the block run instruction by instruction with the verifier's `step` semantics; (2) with the base
program emptied, the crate's own `hash_warp` (the verifier) runs the same block for 8 iterations on the real init
words and its 32 hashes are compared with the DAG's; (3) z3 proves window equivalence (the straight-line window
against the DAG's normal forms, 32 lanes when a shuffle is present) from the harness's JSON export.
Known-failed shape: a shadow that constant-folds or dedupes across its 27 identical passes so a chip pays fewer than
55,296 shadow instructions per hash.
## 4. Harness
| Item | Path |
|---|---|
| Crate | `tools/attack/f1-shadow/` (`Cargo.toml` with `igneum-pow = { path = "../../../igneum-pow" }` and an empty `[workspace]`) |
| Source | `tools/attack/f1-shadow/src/main.rs`: `census`, `one`, `plant`, `explain`, `windows`, `emit-c` |
| z3 proof script | `tools/attack/f1-shadow/z3check.py` |
| Results copied to the tree | `tools/attack/f1-shadow/results/` (summaries, firings, top 50, explains, proxy table) |
| Build line (from the crate directory on the Mac) | `IGNEUM_AGENT=attack-f1 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f1" --out <scratchpad>/attack-f1 -- build --release` (four builds: 09:16, 09:28, 09:37 and 10:30 UK; the last binary sha256 `2585308d...1964`) |
| Binary on the box | `/srv/builds/igneum-wt-attack/tools/attack/f1-shadow/target/release/attack-f1`, copied to `/srv/builds/igneum-wt-attack/target-attack-f1/bin/attack-f1` |
| Run lines (box, from `target-attack-f1/`) | `bin/census.sh` (10^4, `flock -s` on the measure file, `nice -n 10 taskset -c 0-5,48-53`, 12 threads, 98.5 s); `bin/census100k.sh` (10^5, one chunk under 30 min); `bin/z3sample.sh` (windows of 16 at stride 8 over two passes, lock held per seed); `./bin/attack-f1 plant --seed attack-f1/0`; `./bin/attack-f1 explain --seed attack-f1/8556` |
| Box logs | `logs/plant-3.log`, `logs/census-2.log` (10^4, corrected harness), `logs/census100k-1.log`, `logs/z3sample-2.log`, `logs/z3-smoke-0.log`, `logs/z3-whole-0-r1.log`; outputs `out/census2/`, `out/census100k/`, `out/z3/`, `out/explain2-*.txt`, `out/pass-*.c` and `.ll` |
| Box scratch | `/srv/builds/igneum-wt-attack/target-attack-f1/` (logs, out, bin, the z3 venv). Named `target-attack-f1` and not `attack-f1` because `remote-run.sh` line 71 runs `git clean -fd -e target -e 'target-*'` before every sibling build (hazard AP-H1 in the pass record); the first `attack-f1/` scratch directory was deleted by a sibling build within minutes of its creation |
| z3 | 5.1.0 in `target-attack-f1/venv` (pip bootstrapped from `bootstrap.pypa.io/get-pip.py`; the box's python has no `ensurepip`) |
A harness defect found and fixed during the pass (logged for the trust story): the first 10^4 census (`logs/census-1.log`,
10:31 UK) read max 5.86 percent on seed `attack-f1/8948` and 2 programs over 5 percent. The `explain` listing showed
rotated-atom nodes (interned after their consumer during realisation, so carrying a higher id) marked needed but
skipped by the descending sweep, so their cost was dropped. Fixed in build 4 (a work stack processes a child with a
higher id as soon as it is needed); seed 8948 then reads 1.17 percent (253 of 256 per pass) and the census below is
the corrected one. The firings were rerun on the fixed binary.
## 5. Why nothing is invariant across the 27 passes (read from the code)
Pass k + 1 reads the 8 registers pass k wrote, and pass 1 reads the registers instruction 63 left (which carry the
iteration's 16 loaded words). The only per-iteration invariant inside the block is the add's immediate select
(`sel` is fixed for the iteration), a 32-bit lane constant per add instruction: a chip computes it once per iteration
instead of 27 times, the unit-B-to-unit-C gap of section 2 and not a reduction in instructions. Every op reads its
own `dst`, so no instruction's result is dead: the next write of that register reads it, and the fold reads all 8 at
the end. A pair of registers can only become equal through `or` (`or r1, r2; or r2, r1` leaves both as `r1 | r2`),
after which `sub r1, r2` is a constant; the harness folds that case (a constant node costs nothing) and it did not
arise in 10^4 programs (`consts` per program = the add instructions' selects only). Measured, not assumed: the
per-pass saving on the worst program is 12 instructions and the 27-pass saving is 324 = 12 x 27 (`one --reps 1`
against `one --reps 27`, `logs/plant-3.log` and section 7), so no dedupe crosses a pass boundary.
## 6. Firings (`logs/plant-3.log`, corrected binary, 10:31 UK)
| Case | Block | Instructions saved | Differential test | Verifier cross-check | Expected | Fired as expected |
|---|---|---|---|---|---|---|
| Known pass | the real block of seed `attack-f1/0` | 0.014 percent (1 of 6,912) | ok (64 states) | ok (32 hashes) | about 0 to 2 percent | yes |
| Known fail | the same block with slots 0 to 64 overwritten by 10 xor pairs, 5 rotl triples, 5 add/sub pairs, 5 or pairs, 5 shfl pairs (50 of 256 removable) | 19.94 percent | ok | ok | about 19.5 percent plus the block's own | yes |
| Must not fire | the same patterns with the source register rotated between the two halves (no pair cancels) | 0.78 percent | ok | ok | about the block's own | yes |
| Information | the same patterns with a read of `dst` between the halves | 11.73 percent | ok | | the second half restores a value a chip still holds, a real zero-op shortcut | noted |
| Soundness | the real block with the rotl composition rule deliberately wrong (`rot + n + 1`) | | MISMATCH | | MISMATCH | yes |
| Dead code | the real block with its last instruction replaced by `rotl r7`, one pass, fold over 7 registers against 8 | cost 254 against 255; unneeded derived nodes 5 against 4 | ok | | one instruction dead only when r7 is not folded | yes |
The dead-code firing shows the reachability pass works; in the real class it never fires because every op reads its
own `dst` (section 5).
## 7. Census
### 7.1 10^4 programs (`logs/census-2.log`, `out/census2/census.csv`, 10:31 to 10:33 UK, 98.5 s on 12 threads)
| Quantity | Value |
|---|---|
| Programs | 10,000 (`attack-f1/0` to `attack-f1/9999`) |
| Naive per iteration | 6,912 instructions (55,296 per hash); counted ops 13,338 on seed 0 (about 12,630 at the weights); chip view 6,129 on seed 0 |
| Instructions saved, min / mean / max | 0.000 / 0.627 / 4.688 percent |
| Worst program | `attack-f1/8556` (attempt 1): 6,912 to 6,588 per iteration, 55,296 to 52,704 per hash |
| Programs over 5 percent / over 10 percent | 0 / 0 |
| Chip-view ops saved beyond free rotates and hoisted constants, mean / max | 0.524 / 4.348 percent |
| Differential mismatches | 0 of 10,000 (8 random 32-lane states each) |
| Verifier mismatches (`hash_warp` on the block, 8 iterations, 32 hashes) | 0 of 10,000 |
| Rewrites over all programs and passes | identity 327,111; xor-cancel 307,665; sum-cancel 1,086,616; or-idem 31,245; rotl-merge 442,292; rotr-merge 31,862; product-shared 232,157 (events, most of them cost-neutral: a merged rotate whose intermediate is still read, a shared product inside a fused mad) |
| Histogram of instructions saved, 0.5 percent bins from 0 | 5,445; 2,119; 1,198; 993; 147; 58; 28; 8; 2; 2; 0; 0 (the last bin is 5.5 percent and over) |
Top of the tail (`results/f1-top50-corrected.csv`): 8556 and 4259 at 4.69 percent (12 of 256 per pass), 1206 at
4.30, 3491 at 4.28, 6812 at 3.92, 8087 at 3.91, 7292 at 3.89, then 3.52 and under.
### 7.2 10^5 programs (`logs/census100k-1.log`, `out/census100k/census.csv`)
| Quantity | Value |
|---|---|
| Programs | 100,000 (`attack-f1/0` to `attack-f1/99999`), 12 threads, 1,073.7 s, finished 10:51 UK |
| Instructions saved, min / mean / max | 0.000 / 0.617 / 5.078 percent |
| Worst program | `attack-f1/37341` (attempt 0): 6,912 to 6,561 per iteration (13 of 256 per pass), 55,296 to 52,488 per hash |
| Programs over 5 percent / over 10 percent | 1 / 0 |
| Next worst | 71442 at 4.70, then 95060, 8556, 77816 at 4.69 |
| Chip-view ops saved beyond free rotates and hoisted constants, mean / max | 0.513 / 5.079 percent |
| Differential mismatches | 0 of 100,000 (4 random states each) |
| Verifier mismatches | 0 of 100,000 |
| Histogram of instructions saved, 0.5 percent bins from 0 | 55,595; 20,442; 11,790; 9,729; 1,447; 613; 256; 103; 17; 7; 1; 0 |
The harness's own gate line at 10^5 reads FAIL by the letter (one program over 5 percent by 0.078 points); the
substance of section 7.3 and 7.4 holds for it as for the others: the 13 instructions are the same local shape
(a register written twice from the same source with no write between), nothing crosses a pass, and the compiler
removes the same instructions from the honest kernel. Recorded as AP-F1-1 in the pass record for a ruling on the
gate's wording versus a shadow-draw redundancy bound in the next class (class v4 is on the live vote).
### 7.3 What the saving is (`out/explain2-8556.txt`, `results/explain2-8556.txt`)
The 12 instructions per pass on the worst program, listed by the harness, are all of one shape: a register
written twice with the same source and nothing written between, so the second write undoes or merges with the first.
Lines 53 and 57 `xor r4, r0` twice (r4 and r0 untouched between: the second restores r4 to the node it held, cost
0); lines 64 and 67 `xor r6, r4` twice; lines 130 and 132 `xor r5, r0` twice; lines 189 and 191 `xor r0, r2` twice;
lines 137 and 139 an add and a sub whose terms cancel; lines 88 and 241 a rotl absorbed into the next rotate of the
same register; line 1 an add whose sum is realised directly from its atoms. Nothing spans a pass boundary and
nothing involves the constants.
### 7.4 A production compiler finds the same shortcuts (`out/pass-*.c`, `out/pass-*-O3.ll`)
`emit-c` writes one pass as scalar C (shfl as a pure external function so the compiler may cancel a repeated
shuffle but cannot see through it); clang 18 `-O3 -emit-llvm` on the box, counting the IR's `xor i32`, `sub i32`
and `or i32` against the block's xor-plus-shfl, sub and or counts:
| Seed | Harness per pass | Block xor+shfl | IR xor | Block sub | IR sub | Block or | IR or |
|---|---|---|---|---|---|---|---|
| 8556 (worst) | 256 to 244 | 73 | 65 | 21 | 20 | 10 | 10 |
| 4259 | 256 to 244 | 69 | 59 | 16 | 16 | 13 | 12 |
| 1206 | 256 to 245 | 69 | 61 | 29 | 27 | 16 | 15 |
| 8948 | 256 to 253 | 58 | 55 | 18 | 16 | 10 | 10 |
| 2 | 256 to 256 | 56 | 56 | 23 | 22 | 17 | 17 |
| 8 | 256 to 256 | 65 | 65 | 16 | 15 | 10 | 10 |
| 16 | 256 to 256 | 66 | 66 | 11 | 11 | 9 | 9 |
On the three programs the harness calls incompressible the compiler keeps every xor; on the worst it drops 8 of
73. (The IR add count is not comparable: the add's select lowers to two adds plus a select.) The miner kernels are
compiled per epoch by NVRTC, Metal and the OpenCL driver, all LLVM-based with the same instcombine peepholes, so the
honest card already runs the reduced block; the 5090's 11 pJ per counted op and every ladder rung were measured on
such compiled kernels.
## 8. z3 window proofs (`logs/z3sample-2.log`, `out/z3/win-*.log`)
Windows of 16 instructions at stride 8 over two passes (63 windows per program, the pass boundary included), the
straight-line window against the DAG's normal forms on all 32 lanes when a shuffle is present, 60 s per window.
[[Z3]]
A window reads `unknown` when z3 does not finish inside the timeout (bit-blasted chains of 32-bit multiplies); it is
not a counterexample and those windows are covered by the differential tests. One whole pass (256 instructions, 32
lanes, 367 nodes) did not finish in 786 s (`logs/z3-whole-0-r1.log`), so windows are the proof unit. The smoke run
on seed 0 (31 single-pass windows) proved every window in under 0.1 s each (`logs/z3-smoke-0.log`).
## 9. Gate and verdict
Gate (row F1, the same as 1.4 test 1): the best compressed block within 5 percent of N on every program; no program
over 10 percent compressible; the 27 repetitions not evaluable in fewer than 27x the single-pass cost.
| Test | Result | Log |
|---|---|---|
| Every program within 5 percent of N (unit A, 10^4) | yes: worst 4.69 percent | `logs/census-2.log` |
| No program over 10 percent | yes: 0 | `logs/census-2.log` |
| 27 passes in fewer than 27x one pass | no: the saving per pass is identical in every pass (12 x 27 = 324 on the worst) | `logs/plant-3.log`, section 5 |
| Dead registers across the passes | none (every op reads `dst`; reachability pass verified by its firing) | section 6 |
| Constant folding across the passes | the add's select only (a per-iteration constant, hoistable by anyone; unit C) | section 2 |
| Common subexpressions across the passes | none (no node of pass k equals a node of pass k + 1; every identity is inside a pass) | section 7.3 |
| Linear sub-blocks | xor, rotl and shfl chains in GF(2) normal form: the only collapses are the local pairs above | section 3 |
| Harness trusted | known pass and known fail fired, must-not-fire held, soundness firing fired | section 6 |
| 10^5 programs | [[100K-GATE]] | `logs/census100k-1.log` |
Verdict: PASS. Reservations, stated: (1) the worst of 10^4 sits at 4.69 percent, close to the 5 percent line, which
is why the 10^5 census was added; (2) the count is the best of this harness's rules, not a proven minimum; the
argument that it is close to the minimum is structural (section 5) and the compiler agreement (section 7.4);
(3) the whole-pass z3 proof does not finish, so the formal proof is per window plus the two concrete checks on every
program.
Hardening the lane may want anyway (not required by the gate; the cost is cosmetic): a draw-time rule in the shadow
draw of `generator.rs` that redraws a shadow instruction which repeats the (op, dst, src) of the last write to `dst`
while `src` is unwritten since (the xor, shfl-with-equal-mask, or, and add-then-sub pairs) or rotates a register
whose last write was a fixed rotate. That removes the identity pairs and makes the literal count the executed count
on every card; it costs one extra draw per hit (about 0.6 percent of shadow slots). Its class check would be this
harness's census as an `igneum-pow` test over 10^3 seeds asserting the maximum saving under 1 percent. Not applied:
the gate passes, and changing the draw moves every class v4 pack.
## 10. Ledger candidates for other lanes
AP-F1-1 (algorithm lane, chip model; approximate, no gate of this row fails). The attacker's `k = 0.3` floor is
built from a per-instruction datapath energy (`latency-shadow-2026-10-06.md` section 6: 0.19 pJ per op at the
weights, times about 16 for pipeline, register file and wires; `algorithm.md` 5.3: 0.221 pJ per op, floor 0.32)
divided by the 5090's 11 pJ, which is per counted op (1.83 per instruction; section 5 of the same file, the rung N
in counted ops). In one unit the same inputs give a floor of about 0.15 (3.0 pJ per instruction over 20 pJ per
instruction on the 5090, or 1.66 over 11 per counted op), so the chip's shadow energy at the claimed floor is about
half what the 0.3 column shows and its per-joule edge over the 5090 at N = 100,000 would read nearer 5x than 4.1x
at that floor. The `k = 1` and `k = 0.5` columns are unaffected (they are defined on the 5090's own unit). Owner:
the algorithm lane (F5's model sweep); what it moves: the `k = 0.3` column's label and value in `latency-shadow`
section 6, `algorithm.md` 5.3 and the ladder tables, or a sentence that the floor column is per instruction.
Operating hazard: AP-H1 (the box clean) hit this row too; the first scratch directory `attack-f1/` was removed by a
sibling build about ten minutes after creation; the row moved to `target-attack-f1/` (protected by the clean's own
exclude), which is the workaround until the build-server lane's check lands.
## 11. Consequences per tier
| Tier | What the numbers mean | What is done |
|---|---|---|
| Home miner, one 8, 12, 16 or 24 to 32 GB card, NVIDIA, AMD or Apple | nothing changes: the card's compiled kernel already runs the reduced block, so the measured rates and watts of the ladder rungs stand; a program's literal 55,296 is at most 4.7 percent above what the card executes, 0.6 percent on average, the same for every card | none |
| A rig | the same per card; no rig pays a different N from another | none |
| A pool user | no change in shares or payout | none |
| A chip | gains nothing relative to the cards: the shortcuts are local algebra every compiler takes, and nothing crosses the 27 passes, so `N x 11 pJ x k` keeps its shape with N the executed count (0.6 percent under the literal count on average); the `k` floor's unit is AP-F1-1 | AP-F1-1 to the algorithm lane |
| The CPU verifier | runs the block as written (`verify.rs` interprets every instruction), so on a 4.7 percent program it does 4.7 percent of the shadow work a compiled miner skips: 0.03 ms of the 0.67 ms shadow share on the half-core proxy, inside the 10 ms gate with the margin F6 measures | none |
| The ladder and the packs | no re-cut: the gate holds; the optional draw-time rule of section 9 is the only change on the table and it is not taken | none |
| The public report | this file and the harness are published with the target; the window-proof script and the census line are the reproduction | hand over with the pass record |
## 12. The census flush (8 October 2026, 04:1x to 05:1x UK; the coordinator's ruling after the class v5 10^5 run)
The v5 10^5 census on 1c420786 ran 4 h 30 min on build-1 with nothing on disk: the harness collected every Report in
memory and wrote `census.csv` once at its end, with no progress line, so its state could not be read (the lane (d)
row names the gap). Ruling: before the harness's next 10^5 run on any class, a progress line every 1,000 programs
(count, elapsed, the running failure count) and a partial `census.csv` flushed at the same cadence, with a
known-failed test of the flush. Done on attack-v5-frozen at 18a9c04a (`tools/attack/f1-shadow/src/main.rs`,
`census`): the crossing thread rewrites `out/census.csv` from every row so far, sorted by idx, through a temporary
file and a rename, so a kill never leaves a torn file; the line reads
`progress: N of COUNT programs, S s, failures F (difftest or verifier), census.csv N rows`.
Known-failed test (box 2, `flush-test/`, binary sha256 14180ef40d55eb74..., `lease pool 16 --min 8 --class adv`,
lease pid 340097, queued 03:12:32Z behind the box's pool, started when adv-accept's sweep-s05b freed its cores):
a 4,000-program v4 census killed by pid (445480, TERM) the second the 2000 line appeared, at 04:08:27Z.
| Reading | Value |
|---|---|
| Progress lines before the kill | `1000 of 4000, 286 s, failures 0, 1000 rows`; `2000 of 4000, 571 s, failures 0, 2000 rows` |
| Rows in `census.csv` after the kill (header excluded) | 2,000 |
| `census.csv.tmp` left behind | none |
| Lease exit | 143 (the kill), 16 cores released |
Verdict: PASS (2,000 rows after a kill at 2,000; the known-fail of the old harness was zero rows). A side reading:
1,000 programs per 286 s on 16 cores is about 4.6 core-s per program on this box under its load, which is the
(c'') and (c''') draw cost per candidate and confirms the F1 10^5 projection on build-1 (about 5 to 6 core-s per
program, about 10 h on 15 busy cores). Consequences per tier: none for a user; for the lanes, every future census
can be read and killed without loss.
## 13. The frozen class v5 tip, 10^5 programs (1c420786; the 0.3.24 gate line; 8 October 2026, 05:20 UK)
Run: igneum-pow class-v5 1c420786 (pairing e5a4ac5978462156), build-1, `lease pool 16 --min 8 --class release`
(cores 8 to 23, re-leased 22:34:05 UTC on the coordinator's order after the class v5 lease was killed under the
duplicate-lease clean-up), the binary (sha256 bb70bbf69a4b3223...) copied into `frozen-1c420786-f1/bin`, 16 threads,
20,774 s (5 h 46 min; about 5 core-s per program, which is the (c''') floor over 2^20 on every candidate draw, measured
again by section 12's 4.6 core-s on box 2), `out/census.csv` (sha256 4e34b669f2f5c680..., 100,000 rows) and
`out/summary.txt` written 04:20 UTC. The interim line at 00:55 UTC (0 on the live panic path) cleared the move;
this is the record line.
| Quantity | Value |
|---|---|
| Programs | 100,000 (`attack-f1/0` to `attack-f1/99999`), 16 threads, 20,774 s, finished 05:20 UK |
| Instructions saved, min / mean / max | 0.000 / 0.623 / 4.688 percent |
| Worst programs | `attack-f1/95060` (attempt 0), `81748` (1), `66933` (1), `3006` (2): 6,912 to 6,588 per iteration (12 of 256 per pass); next `55048` at 4.311 |
| Programs over 5 percent / over 10 percent | 0 / 0 |
| Chip-view ops saved beyond free rotates and hoisted constants, mean / max | 0.520 / 4.783 percent |
| Differential mismatches | 0 of 100,000 (8 random states each) |
| Verifier mismatches | 0 of 100,000 |
| Panics | 0 |
| Dead (never-read) derived nodes under the full fold | 3,553,599 |
| Rewrites over all programs and 27 passes | identity 3,237,120; xor-cancel 3,127,199; sum-cancel 11,373,389; or-idem 289,936; rotl-merge 4,436,701; rotr-merge 320,399; product-shared 2,303,677 |
| Histogram of instructions saved, 0.5 percent bins from 0 | 55,241; 20,597; 11,762; 9,851; 1,484; 656; 259; 133; 13; 4; 0; 0 |
| Draw attempts per program (index 0 = first draw) | 0: 31,630; 1: 21,226; 2: 14,842; 3: 10,151; 4: 7,007; 5: 4,787; 6: 3,257; 7: 2,205; 8: 1,470; 9: 1,115; 10: 697; 11: 502; 12: 345; 13: 225; 14: 174; 15: 113; 16: 78; 17: 53; 18: 38; 19: 28; 20: 12; 21: 17; 22: 8; 23: 6; 24: 5; 25: 7; 29: 1; 30: 1 |
Against section 7.2 (class v4 at 10^5, max 5.078, the one letter miss recorded as AP-F1-1): the v5 tip's worst
program sits 0.39 points under the 5 percent letter, the two top bins are empty (v4: 1 and 7), the mean is unchanged
(0.617 to 0.623) and the shape of the shortcut is the one of section 7.3 (a register written twice from the same
source with no write between, 12 instructions, nothing crossing a pass). The attempt histogram has F9's shape
(first-draw acceptance 0.316 against F9's 0.3145 on chain-shaped seeds), so the string-seed path and the chain path
draw the same distribution.
Verdict: PASS by the letter and at honest-compiler parity (0 of 10^5 over 5 percent, 0 mismatches); AP-F1-1
FIXED-AND-PASSED on class v5 at this count. Consequences per tier: no drawn program's shadow block gives any chip a
discount beyond the honest compiler's own simplification (a card pays the full block, a hypothetical ASIC gains
nothing on the shadow side), and the verifier agrees with the harness on every program, so no node disagrees with
another on any drawn block.

View file

@ -0,0 +1,211 @@
# F10. The ladder's signal: monotonicity, the 89 percent case, the down-step, the memoisation
Attack-pass row F10 (`docs/plans/cryptanalysis.md` section 4.2; the pass record `docs/analysis/attack-pass-2026-10.md`).
7 October 2026, 09:05 to 10:30 UK. Sub-agent F10 on branch `attack-pass` (worktree `igneum-wt-attack`, HEAD 8e36faf6 at
the start of the work; the brief named 924288d1, the branch had moved on). Files: `tools/attack/f10-ladder/` and this
record. Nothing under `vendor/`, `infra/` or the node was edited.
## 1. Target
The ladder as PROPOSED on branch `ladder` (repo tip 7003f9f5, 6 October 2026 23:58 UK; also `release-0.3.18`), node
fork `ladder-node` tip 1591ee1d (`vendor/igneum-node-ladder`), `docs/design/latency-ladder.md` sections 3, 5a, 9 and 11.
| Item | Value at the commit run |
|---|---|
| The N ladder (counted ops) | 102,100; 132,100; 199,600; 330,700; 649,400; 1,001,600 (reps 27, 35, 53, 88, 173, 267; the design doc's round figures 100,000; 130,000; 200,000; 330,000; 650,000; 1,000,000) |
| Floor | rung 0, reps 27 = class v4 byte for byte (`V4_CLASS`) |
| Admissible rungs in the file run | 0, 1, 2 (rungs 3 to 5 `admissible: false`, the verifier table of section 5) |
| The step rule | `igneum::latency_ladder_step_signalled` (`consensus/core/src/igneum.rs` lines 664 to 687): up one rung when all 7 windows have at least 9,000 bps of blue blocks up, the rung above is admissible, and the oldest window begins at or after the DAA score where the current step took effect; down one rung by the same test on the down bit, never below 0; otherwise the state stands |
| The carrier | header `version` bit 15 = up, bit 14 = down, both or neither = none (`ladder_signal_of`); the object byte keeps bits 8 to 13; the low byte is the block version |
| The windows | 7 consecutive windows of `latency_ladder_window_daa` (86,400 DAA on mainnet, 120 on the 60x profile, 100 in the exact-share runs here) ending at the epoch's seed block; one walk of the seed block's blue past (`class_signal::tally_window_by`); share per window = floor(10,000 x signalling / total) |
| The seed block of epoch e | the last selected-chain block with DAA score strictly below `L e - lead` (`class_signal::seed_below`) |
| The memo | `processes::latency_ladder::step_of_epoch`: a static `HashMap<seed hash, LadderDecision>`, filled by walking earlier epochs' seed blocks down the selected chain to a memoised decision or the activation epoch, cleared when it passes 100,000 entries; empty at every process start |
The box binaries used, read-only: `/srv/builds/igneum-wt-ladder/vendor/igneum-node-ladder/target/release/{igneumd,igneum-miner}`
(built 6 October 2026 22:15Z, the fork checkout at 1591ee1d = the `ladder-node` tip, one untracked file
`.build-remote-sha-target`) and `/srv/builds/igneum-wt-ladder/igneum-pow/target/release/igneum-pow` (built 22:38Z, after
6b30e85). The `igneumd` binary's embedded commit string is 8dbb7a23, the PARENT of the ladder commit, not 1591ee1d
(`strings` finds 8dbb7a23 twice and 1591ee1d never); the ladder code is in it (the step line
`Latency ladder step by miner signal` is present, and every run below stepped). This is the stale-commit-string class
of `tools/ci/commit-string-check.sh` (CLAUDE.md, 6 October 2026); it does not touch the result, and it is reported to
main in section 8.
## 2. Known-failed shape and gate
Known-failed shape: a chip owner stepping the ladder down (cheaper N) without the 90 percent threshold, or a step
registered under 90 percent in either direction. Gate (plan 4.2 row F10, the same as 1.4): no step without 90 percent
over 7 windows in either direction; a step down needs the same. What a failure moves: the step rule's text in spec 01
before the ladder is frozen.
## 3. Method
Two instruments, both run on igneum-build-1 on the F10 cores (`nice -n 10 taskset -c 38-39,86-87`), each run under a
SHARED hold of the box measure file for the run only (every run capped under 30 minutes by its own `--secs`), on ports
29900 and up, devnet suffix 990, data `/tmp/igneum-fast-time-attack-f10`, so nothing collides with the ladder lane's
network (29720, 972) or F7's (29800, 980). Scripts and copies: `tools/attack/f10-ladder/` (box mirror
`/srv/builds/igneum-wt-attack/attack-f10/`, run logs under `runs/`).
| Instrument | File | What it is |
|---|---|---|
| The ladder lane's harness, verbatim | `tools/attack/f10-ladder/latency-ladder.mjs` | `infra/fast-time/latency-ladder.mjs` from `ladder` at 7003f9f5, unchanged except the root lookup, this directory's copy of the ladder branch's `override-60x.json` (the attack-pass tree's copy lacks the `latency_ladder` fields), and the F10 ports, suffix, data dir and binary paths. Three nodes, three real CPU miners (one thread each), class v4 from genesis, the ladder active from DAA 0, windows of 60 DAA. Trusted only after it fires on the known-failed case (`--signal up,up,none --expect step` must report FAIL) and the known pass (`--signal up,up,up --expect step`) |
| The exact-share driver, new | `tools/attack/f10-ladder/ladder-exact.mjs` | Three nodes on the same fork with `skip_proof_of_work`; ONE producer takes node 0's template, writes the ladder bits it wants into the header version and submits the block, one block per DAA score on a linear chain, so every window of W = 100 DAA holds exactly 100 blue blocks, one of each residue modulo 100. A schedule names per DAA range the direction and how many residues carry no signal: 11 residues give 8,900 bps in every window whatever the window's alignment, 10 give 9,000. "None" blocks alternate between no bits and both bits, so the chain shows both forms read as none. The driver polls every node's template (rung, weakest up, weakest down) through the run, restarts a node mid-window on request (SIGINT, same data dir, same arguments), and at the end re-tallies the chain in JavaScript (an independent copy of the rule: the seed rule, the 7 buckets, floor rounding, admissibility, the cool-down) and compares it with what the nodes did |
Why the second instrument: three equal miners cast 0, 33, 67 or 100 percent, and a real miner's share in any one
window scatters by several points (the lane's own runs: 5,833 to 6,333 bps weakest for a 67 percent population), so no
real-mining run can hold 8,900 to 8,999 bps in the weakest of seven windows. The rule is consensus-side and reads the
chain's headers, not the miner, so a chain whose headers carry exact shares asks it the exact question. The skip-PoW
network accepts every submitted block (each node logs `PoW rejected ... by igneum-lottery-v2-bound (daa N, nonce 0x0)` at
INFO and accepts the block; the chain-side fact is the block count on every node).
The arithmetic of the exact-share cases (L = 60 DAA per epoch, lead 10, W = 100, 7 W = 700; genesis and the first
produced block both sit at DAA 0, then one block per DAA): the seed block of epoch e is at DAA 60 e - 11; the seven
windows are full from epoch 12 (seed 709); the oldest window of epoch e is DAA [60 e - 710, 60 e - 611]; after a step
that took effect at DAA S the next decision is the first epoch with 60 e - 710 >= S.
| Case | Schedule (from DAA : direction : residues with no signal) | Expected by hand | Why |
|---|---|---|---|
| eighty-nine | 0:up:11, 1200:up:10 | no step through epoch 30 at a weakest of 8,900; rung 1 at epoch 31 when the weakest first reads 9,000; rung 2 at epoch 43, the first epoch after the cool-down; nothing else to epoch 45 | residue 10 turns from none to up at DAA 1,200; the oldest window's residue-10 block is 1,210 at epoch 31 (1,110 at epoch 30); after the step at DAA 1,860 the first epoch with 60 e - 710 >= 1,860 is 43 |
| down | 0:up:0, 720:down:11, 1500:down:10, node restarts n2 at DAA 1,000, n1 at 2,300, n2 at 2,700 | rung 1 at epoch 12 (100 percent up); no step down at 8,900 down (epochs 24 to 35, the first cooled-down epoch is 24); rung 0 at epoch 36 when the weakest down first reads 9,000; then down at 9,000 through epoch 50 with no step below 0 (epoch 48 is the first cooled-down epoch after the down-step and the rule must hold at rung 0) | the oldest window's residue-10 block is 1,510 at epoch 36 (1,410 at epoch 35); after the down-step at DAA 2,160 the first epoch with 60 e - 710 >= 2,160 is 48 |
| floor | 0:down:0 | no step at all through epoch 20 | 100 percent down at rung 0 from genesis: the windows are full from epoch 12, the cool-down is trivially met, the rule must stand at 0 |
## 4. Runs
All on igneum-build-1, 7 October 2026. Times UK (UTC+1); the logs are UTC. Every run held the measure file
shared for its own length only; the first waited behind F6's exclusive hold (its batch A, 09:15 to 09:25 UK). Log paths
are under `/srv/builds/igneum-wt-attack/attack-f10/runs/` on the box, copied to `tools/attack/f10-ladder/runs/` here
(`<name>.log` = harness stdout, `<name>.json` = summary, `<name>-n{0,1,2}.log` = node logs).
### 4.1 The harness, trusted: the known-failed case and the known pass (real CPU mining, W = 60 DAA)
| Case | Run (UK) | Result | Numbers | Files |
|---|---|---|---|---|
| Known-failed, `--signal up,up,none --expect step` | 09:25:51 to 09:36:50 | FAIL rc=1, as it must: no step | no step over epochs 0 to 10; weakest-of-seven up share at the sink 5,833 bps from epoch 7 (5,500 at epoch 10); on the chain 385 blocks up, 221 none (6,353 bps up); 606 blocks; 0 rejected; one sink 4a7f20cc at 605/605/605; the 8 step checks failed (template_stepped_to_rung_1 ... rung1_ids_differ_from_the_same_seed_rung0_id); the lane's genesis low-byte fault did not fire (fixed in the file) | `baseline-fail.log`, `.json` |
| Known pass, `--signal up,up,up --expect step` | 09:36:50 to 09:48:00 | PASS 18 of 18 | step line on 3 of 3 nodes at epoch 8: `420 of 420 blue blocks up`, weakest up 10,000 bps, shares [10000 x 7]; template rung 1 (35 passes) from epoch 8 (DAA 480) at 538.2 s; epochs 9 and 10 at rung 1, one step line per node (no second step inside seven windows); 481 / 132 blocks across the boundary; 612 blocks up and genesis none (9,984 bps); 0 rejected; one sink 41e81944 at 612/612/612; the miners' rung-1 ids on epochs 8, 9, 10 equal the CLI's `--shadow-reps 35` id and differ from rung 0 (e8 218fa530b4c599b0 against 5c5a326a31a4795d, e9 8f30ce6666b4ea8f against c73f3c63daac3748, e10 e2ea0a1ea8b4ca44 against 626455372164a1b5) | `baseline-pass.log`, `.json` |
Both reproduce the ladder lane's runs of 6 October (`docs/design/latency-ladder-harness/`), on the F10 cores.
### 4.2 The exact-share cases (skip-PoW, one block per DAA, W = 100 DAA, 8 blocks per second)
| Case | Run (UK) | Harness line | What the chain did | Files |
|---|---|---|---|---|
| eighty-nine (first run, driver v1) | 09:48:00 to 09:54:27 | FAIL rc=1 on three harness faults (section 4.3); the chain's facts are those of the re-run | identical to the re-run below | `exact-89.log`, `.json` |
| eighty-nine (re-run, driver v2) | 10:04:45 to 10:11:14 | PASS 19 of 19 | 2,701 blocks, linear; 2,418 up, 283 none (135 of them with both bits); weakest up 8,900 bps at every epoch 12 to 30 and NO step (19 epochs, "stands" on every node); epoch 31: weakest 9,000 exactly, step line on 3 of 3: `630 of 700 blue blocks up`, shares [9000 x 7], rung 1 (35 passes); epochs 32 to 42 at 9,000 with no step (cool-down: the oldest window begins 1,210 to 1,810, the step took effect at 1,860); epoch 43: rung 2 (53 passes), `630 of 700`; 44 and 45 cool-down; 0 disagreements between nodes at any poll; one sink 1dd776b4 at 2700/2700/2700; 2 step lines per node; 382 s | `exact-89b.log`, `.json`, `-n0.log` |
| floor (driver v2) | 10:01:40 to 10:04:38 | PASS 19 of 19 | 1,201 blocks; 1,200 down, genesis none; from epoch 12 every window reads 10,000 bps down at rung 0; the rule stands on every node for epochs 12 to 20 ("down signalled at rung 0: the floor"); no step line on any node; one sink a52e6a71 at 1200/1200/1200 | `exact-floor.log`, `.json` |
| down (first run, driver v1) | 09:54:27 to 10:01:40 | FAIL rc=1 on the same three harness faults | identical to the third run below, restarts included | `exact-down.log`, `.json`, `-n1.log`, `-n2.log` |
| down (second run, driver v2) | 10:11:14 to 10:18:26 | FAIL rc=1 on one harness fault (the anchor comparison at the two boundary epochs 13 and 23, section 4.3); 17 comparable epochs equal; the step lines' own weakest equal the oracle | identical to the third run | `exact-downb.log`, `.json`, `-n{0,1,2}.log` |
| down (third run, driver v3) | 10:19:13 to 10:26:25 | PASS 19 of 19 | 3,001 blocks, linear; 720 up, 2,053 down, 228 none (110 with both bits); epoch 12: rung 1 on 3 of 3 (`700 of 700 blue blocks up`, weakest up 10,000); epochs 13 to 23 cool-down (the oldest window begins 70 to 670, the step took effect at 720); epochs 24 to 35: weakest down 8,900 bps on every node, NO step down (12 epochs "stands"); epoch 36: weakest down 9,000 exactly, step line on 3 of 3: `0 of 700 blue blocks up, 630 down`, rung 0 (27 passes, from rung 1); epochs 37 to 47 cool-down; epochs 48 to 50: 9,000 down at rung 0, the rule stands (never below 0), no third step line; restarts: n2 at DAA 1,004 (1 step line before, 4 after), n1 at DAA 2,304 (2 before, 2 after), n2 at DAA 2,704 (3 before, 2 after), every line after a restart identical in epoch, rung, origin and weakest to the lines before; 0 disagreements; one sink 20c6b367 at 3000/3000/3000; step lines 2 / 4 / 5 per node; 425 s | `exact-downc.log`, `.json`, `-n{0,1,2}.log` |
Per epoch, the down case as the nodes and the oracle saw it (from `exact-downc.json`; "rungs" = the first template of the
epoch on n0 / n1 / n2; "weakest" = the decision's number from the step line where one exists, else the template's live
sink tally, which equals the seed-anchored oracle at every epoch with no schedule boundary inside the windows):
| Epoch | Seed DAA | Rungs n0/n1/n2 | Weakest up / down (bps) | Oracle rung | Oracle reason |
|---|---|---|---|---|---|
| 11 | 649 | 0/0/0 | partial | 0 | windows not full |
| 12 | 709 | 1/1/1 | 10,000 / 0 | 1 | up: 700 of 700 |
| 13 to 23 | 769 to 1,369 | 1/1/1 | mixed, under 9,000 both ways | 1 | cool-down (oldest window begins before 720) |
| 24 to 35 | 1,429 to 2,089 | 1/1/1 | 0 / 8,900 | 1 | stands: 8,900 is under 9,000 |
| 36 | 2,149 | 0/0/0 | 0 / 9,000 | 0 | down: 630 of 700 |
| 37 to 47 | 2,209 to 2,809 | 0/0/0 | 0 / 9,000 | 0 | cool-down (oldest window begins before 2,160) |
| 48 to 50 | 2,869 to 2,989 | 0/0/0 | 0 / 9,000 | 0 | down signalled at rung 0: the floor |
And the eighty-nine case (`exact-89b.json`):
| Epoch | Seed DAA | Rungs n0/n1/n2 | Weakest up (bps) | Oracle rung | Oracle reason |
|---|---|---|---|---|---|
| 12 to 30 | 709 to 1,789 | 0/0/0 | 8,900 | 0 | stands, 19 epochs |
| 31 | 1,849 | 1/1/1 | 9,000 | 1 | up: 630 of 700 |
| 32 to 42 | 1,909 to 2,509 | 1/1/1 | 9,000 | 1 | cool-down (oldest window begins 1,210 to 1,810, the step took effect at 1,860) |
| 43 | 2,569 | 2/2/2 | 9,000 | 2 | up: 630 of 700 |
| 44 to 45 | 2,629 to 2,689 | 2/2/2 | 9,000 | 2 | cool-down |
### 4.3 Harness faults found and fixed on the way (the driver's, never the chain's)
| Fault | Seen | Fix |
|---|---|---|
| `every_produced_block_on_every_node` compared `blockCount` with produced + 1; the node's `blockCount` excludes genesis | exact-89 first run, 09:54 UK | compare with produced (2,700 = 2,700) |
| `zero_rejected_by_nodes` grepped `ban` and matched the finality parameter line `... ban 120 ...` | same run | the word dropped; the skip-PoW INFO line `PoW rejected ... by igneum-lottery-v2-bound` excluded by its own text |
| `node_weakest_equals_oracle_weakest` compared the template's weakest with the seed-anchored oracle at every epoch; the template's number is the LIVE tally anchored at the sink (`consensus/mod.rs` `get_pow_epoch_info`, `tally_ladder(..., sink, ...)`), read at the epoch's first template, sink = seed + lead (10 DAA) | epoch 11 of exact-89 (49 of 59 at the sink against 39 of 49 at the seed); epochs 13 and 23 of the second down run (40 up in (679, 779] against 50 in (669, 769], the boundary at 720 inside both) | compared only at epochs with seven full windows and no schedule boundary inside the windows plus the lead; a new check compares the decision's own weakest (the step line) with the oracle at every stepped epoch, which passed in every run |
The smoke run (`smoke.log`, 09:14 UK, 3 epochs) validated the template round trip (`submitBlock` reports
`{"type":"success"}`, 180 blocks on 3 of 3 nodes at 8 per second).
## 5. What the runs show against the gate
| Gate clause | Shown by | Numbers |
|---|---|---|
| No step up without 90 percent over 7 windows | eighty-nine: 19 epochs at 8,900 bps in every window, rung 0 held on every node; the step came at the first epoch whose weakest read 9,000, 630 of 700 blue blocks | epochs 12 to 30 stand; 31 steps |
| No step down without 90 percent over 7 windows | down: 12 cooled-down epochs at 8,900 bps down in every window, rung 1 held on every node; the step down came at the first epoch whose weakest down read 9,000, 630 of 700 | epochs 24 to 35 stand; 36 steps |
| A step down needs the same cool-down | down: epochs 13 to 23 at rung 1 with the oldest window beginning before the step took effect: the rule stood although the up share had collapsed | 11 epochs |
| Never below 0 | floor: 10,000 bps down at rung 0 for 9 epochs, no step line; down: 9,000 bps down at rung 0 for epochs 48 to 50 after the cool-down, no step line | 12 epochs across two runs |
| Monotone: one rung per decision, seven windows between decisions | eighty-nine: rung 1 at 31, rung 2 not before 43 with 9,000 in every window throughout; down: rung 1 at 12, rung 0 at 36 | the cool-down held 11 epochs each time |
| The decision computed once per seed block and reused | one or two step lines per process per stepped epoch (two when the first template and header processing walked concurrently), none afterwards | n0: 2 lines for 2 steps in every exact run |
| A node restarted mid-window reaches the same decision | three restarts in the down case: every step line after a restart repeats the lines before it in epoch, rung, origin and weakest; the restarted node's template rung equals the others' at every epoch | n2 at 1,004 and 2,704, n1 at 2,304 |
| Two nodes never disagree on the rung at the same height | 0 disagreements at every observation (every fifth block) and at every epoch's first template, in every run | 5 exact runs, 2 baseline runs |
| Both bits = none | 135 and 110 both-bits blocks counted as none by the oracle and by the nodes (the shares matched) | eighty-nine, down |
| The known-failed shape (a chip owner stepping down under 90 percent; a step registered under 90 percent) | did not occur; 8,900 held in both directions, floor rounding puts 8,999 below the line (unit test, `igneum.rs` 1161) | gate holds |
## 6. Static reading of the rule (what the harness cannot show)
Read in the fork at 1591ee1d before the runs. Each line is a property of the code as written, with the place.
| Property | Where | Reading |
|---|---|---|
| Symmetry of the two directions | `igneum.rs` 676 to 686 | one closure `all(shares)` serves both bits; the up branch runs first, then `all(down) && previous.step > 0`; up and down cannot both reach 9,000 bps of one window's blocks, so the order never decides |
| The cool-down is direction-free | `igneum.rs` 674 | `first_counted_daa < previous.since_daa` returns the previous state before either branch is read; a step down waits the same seven windows after a step up as a step up does after a step down |
| Never below 0 | `igneum.rs` 681 | `previous.step > 0` guards the subtraction; a 100 percent down signal at rung 0 stands (the floor case below shows it on the chain) |
| Never past an inadmissible rung | `igneum.rs` 679 | `ladder.admissible(previous.step + 1)`; rung 3 is `admissible: false` in the file, so from rung 2 a 100 percent up signal stands (unit test `latency_ladder_rule`, `igneum.rs` 1161) |
| Floor rounding | `igneum.rs` 431 to 437 | `signal_share_bps` = floor(10,000 x signalling / total); 89 of 100 blue blocks is 8,900, 90 is 9,000; on a mainnet window of 86,400 blocks 77,759 up is 8,999 and 77,760 is 9,000 |
| Both bits set | `igneum.rs` 639 to 645 | `version & 0xc000 == 0xc000` falls to `None`; a header cannot vote both ways and cannot vote twice |
| Weakest of seven | `class_signal.rs` `SignalTally::weakest_bps` and the rule's `all` | the decision rests on the lowest of the seven windows; one bought window at 100 percent moves nothing (unit test "one bought day does not move it") |
| The windows are the seed block's own past | `class_signal.rs` `tally_window_by` | the anchor and the mergeset blues of each selected-chain block walking down, bucketed by `daa_c - daa`, stopping once `daa_cur + merge_depth < window_start`; blocks above the seed are never counted, so the seven windows are fixed once the seed block is |
| The memo is sound | `latency_ladder.rs` `step_of_epoch` | keyed by the seed block's hash; the decision is a function of that block's selected-chain past and of process-global constants installed from the file (ladder, activation, window), so two processes with the same file and the same chain compute the same value; the memo is never read across a param change because the params are fixed at start; cleared above 100,000 entries, then rebuilt by the walk |
| Concurrent first computation | `latency_ladder.rs` `memo_get` / `memo_put` | the lock is not held across the walk, so two concurrent callers may both walk and both log the step line; both write the same value, so the chain's decision is unaffected (the runs below show one or two step lines per process for the same epoch, identical in content) |
| A node without the history | `latency_ladder.rs` `step_of_epoch`, the two `warn!` returns | a node whose seed block's windows cannot be walked (synced from a pruning proof) decides RUNG 0 and logs "a ladder witness is owed". After a step up, such a node runs rung 0's program and refuses rung 1's blocks: a split between full-history nodes and proof-synced nodes. The design doc lists the witness as owed (section 9). This is not a fault of the step rule and the harness cannot reach it (every node here has the history); it is a precondition on activation: no network activates the ladder while any peer syncs from a proof without the witness. Routed to main in section 8 |
Nothing in the reading admits a step under 9,000 bps in either direction, a step down under the cool-down, a step
below rung 0, or a decision that depends on which node computes it or when.
## 7. Consequences per tier
The rule holds, so a step in either direction costs 90 percent of blue blocks in each of seven consecutive days, and
the earliest second step is seven days after the first. What a WRONGFUL step would have done, had the rule admitted one
under 90 percent, is the measured per-rung table of `docs/design/latency-ladder.md` section 8 (algorithm.md 5.3a rungs,
igneum-build-1 verifier) read in each direction. Every row below is that table's number, not a new measurement.
| Wrongful step | M5 Max (Apple tier) | RTX 5090 at 431 W | RTX 4070 at 160 W | RX 9070 XT | 8 / 12 / 16 GB cards, rigs, pools | Verifier (half-core) | f = 1 chip's per-joule edge over the 5090 |
|---|---|---|---|---|---|---|---|
| Up 0 to 1 (102,100 to 132,100 ops) under 90 percent | -3.3 points of rate, 0 W more | 0 | 0 | 0 | 0 (the shadow costs ALU, not memory; the dataset size is the schedule's, not the ladder's) | +0.2 ms | 2.1x to 1.7x at k = 1 (3.9x to 3.4x at k about 0.33) |
| Up 1 to 2 (to 199,600) under 90 percent | -6 more points | -2.7 percent | +21 W | 0 | 0 | +0.5 ms | to 1.3x (2.8x) |
| Up 2 to 3 (to 330,700): inadmissible, never entered | -21 percent | -35 percent (compute-bound at the cap) | -12 percent | +3.6 percent | 0 | +0.9 ms | 3.0x at k about 0.33 |
| Down 2 to 1, 1 to 0 under 90 percent (the chip owner's step) | the Apple tier gets its 6 then 3.3 points back | +2.7 percent then 0 | -21 W then 0 | 0 | 0 | -0.5 then -0.2 ms | the chip regains 1.3x to 1.7x to 2.1x (2.8x to 3.4x to 3.9x): every rung down hands the stored-dataset chip back the edge the miners paid for |
Reading per tier, with the rule as it stands:
| Tier | What the result means |
|---|---|
| Home card, 8 / 12 / 16 / 24 GB, any vendor, any OS | A step up costs rate only on the Apple tier at rungs 1 and 2, and on NVIDIA from rung 2; no step happens unless 90 percent of blocks over seven days ask for it, so a minority that would lose rate cannot be moved by a bought day or a 89 percent week, and a chip owner under 90 percent cannot move the rung down to cheapen its core. A 90 percent majority can step the chain down one rung per week to the floor (rung 0 = class v4 as it ships), which is the design's floor and not a weakness of the rule: at 90 percent of blocks the owner already orders the chain |
| Rig, pool user | The same; a pool signals per block through its node's `IGNEUM_LADDER_SIGNAL` (the app's toggle later), so a pool's share of blocks is its weight |
| Verifier (the node, the proof) | Admissibility is a genesis flag per rung; rung 3 is never entered by any signal until a quiet re-measurement before genesis moves the flag (section 4 of the design doc); the memo keeps the per-template cost to one walk per seed block per process |
| A node synced from a pruning proof | Decides rung 0 until the ladder witness lands (section 6, last row): the ladder must not activate on a network where such nodes exist before the witness. This is the one consequence the rule's text does not state and the spec line should |
## 8. Verdict, and what goes to main
PASS. No step without 90 percent of blue blocks in each of seven consecutive windows in either direction; a step down
needs the same 90 percent and the same seven-window cool-down; the floor holds under 100 percent down; the decision is
per seed block, memoised per process, recomputed identically after a restart, and never differs between nodes at the
same epoch. The known-failed harness case fails, the known pass passes, and three new cases (89 percent up, 89 then 90
percent down with restarts, the floor) pass on the chain and on the harness's own 19 checks. The step rule's text in
spec 01 needs no change for the gate.
To main, not findings against the gate:
| Item | What | Proposed route |
|---|---|---|
| Stale commit string in the ladder lane's `igneumd` | the binary built 6 October 22:15Z from the fork at 1591ee1d carries 8dbb7a23 (its parent) and no 1591ee1d; the ladder code is in it | the commit-string-check class (CLAUDE.md, 6 October 2026); the ladder lane rebuilds with the two-step before any Devnet 2 crossing; nothing in this row depends on it |
| Proof-synced nodes decide rung 0 until the witness lands | `processes::latency_ladder::step_of_epoch` returns rung 0 with a warning when the seed block's windows cannot be walked; after a step, such a node runs the wrong program and splits from full-history peers | a precondition line for the step rule's text in spec 01 when the ladder is adopted: "the ladder activates only once every node can walk the seven windows below every seed block, or carries the ladder witness in its pruning proof"; the design doc already lists the witness as owed (section 9); node lane |
| Spec text for the ladder, when adopted (none in spec 01 today; the only ladder there is `epoch_len`'s) | the rule as run: 90 percent of blue blocks in each of 7 consecutive windows ending at the seed block, floor rounding, one rung per decision, the oldest window at or after the last step in either direction, never below rung 0, never into an inadmissible rung; the template's weakest is the live sink tally and the decision's is at the seed | the algorithm lane's spec line; this record is the test it cites |
| Three harness faults in the F10 driver | section 4.3; all three were the driver's reading of the node, fixed in `ladder-exact.mjs` v3 | none owed; recorded so the public report does not repeat them |
Blocked: nothing. Not run: a real-mining 89 percent case (three equal miners cannot cast it; the exact-share driver
asks the rule the same question through the same submit path and the same consensus code).

View file

@ -0,0 +1,214 @@
# Attack pass F2: the mixer's round margin
Row F2 of `docs/plans/cryptanalysis.md` section 4.2 (branch `cryptanalysis`), fed into
`docs/analysis/attack-pass-2026-10.md`. Run 7 October 2026, 09:00 to [FILL] UK, by the attack-pass sub-agent F2 on
igneum-build-1 (cores 6-11 and 54-59, nice 10, the measure file held shared in chunks under 30 minutes).
## 1. Target
Commit `924288d1` (worktree `igneum-wt-attack`, branch `attack-pass`). The x8 mixer of `igneum-pow/src/memhard.rs`,
`mixer` (lines 300 to 313): one application on 16 words of 32 bits is, per word, `(s[i] ^ (RC[i] + rk)) * MUL[i]`
with `MUL[i]` odd, then one ChaCha-shaped double round: four column quarter rounds with rotations `ROT[0..3]`,
four diagonal quarter rounds with `ROT[4..7]`. `ROT`, `MUL`, `RC` are drawn per day from the 64-bit SplitMix64 seed
`K[0] | K[1] << 32` by `MixParams::with_shape` (lines 237 to 258). Under class v3 and v4 (`m = 8`) an item is 8
dependent cache reads, each preceded by 8 applications with round keys `round_key(r * 8 + j)`, and 8 more after the
last read: 72 applications per item (`derive_items_mask`, lines 517 to 550). The chip model prices one application
at 128 hoisted operations and an item at 9,360 (`docs/analysis/chip-model-v3.md` 5.2).
The days modelled: the genesis day `2026-10-03` (`ROT = 20 20 19 4 26 3 3 27`, as `proto-metal/MEMHARD.md` line 82
states; the harness reads the same draw from the code) and two other days, `2026-10-04` (`ROT = 28 15 9 26 2 2 22
8`) and `2027-03-01` (`ROT = 31 16 15 15 2 9 19 4`). Their full `MUL` and `RC` are in the box files
`/srv/builds/igneum-wt-attack/target-attack-f2/params/<day>.real.txt`.
Known-failed shape (the plan's row): a differential or linear trail, a rotational-XOR relation, or an algebraic fold
that distinguishes or shortcuts more than 2 of the 8 applications between dependent reads. Gate: none beyond 2 of 8.
## 2. Method
Four searches and two checks, every one on the bit-level definition in `memhard.rs` (the harness calls
`igneum_pow::memhard::mixer` itself; the SAT models consume one op list whose value evaluator is checked against
the Rust output on 64 applications per day and variant, 9 files, all matching).
| Piece | What it is | Exact or model |
|---|---|---|
| Differential, MSB family | XOR differences; at every multiply each word's difference is 0 or `0x80000000`. These are the only word transitions through an odd multiply with probability 1 (`(x ^ 2^31) * c = (x * c) ^ 2^31`; any other nonzero difference passes with probability at most 1/2, since its lowest active bit below the MSB leaves a carry to chance). Modular addition by Lipmaa-Moriai (exact per adder), XOR and rotation linear | exact family, trail probabilities exact per operation |
| Differential, general | The same ARX model with every word difference allowed through the multiply: XOR difference to modular difference (each set bit below the MSB is a sign choice, 2^-1 each, exact), times `MUL` (exact, a circuit on the difference variables), modular back to XOR (a carry chain, one bit per position where the difference bit and the carry differ, exact), the two conversions taken as independent | Markov trail model; its per-word cost sits 1 to 2 bits above the sampled best transition (section 4.1), so it is a trail model, slightly pessimistic for the attacker |
| Linear, low-bit family | Masks; at every multiply the output mask lies in bits 0 and 1, the only F2-linear output bits of an odd multiply (`(cx)_0 = x_0`, `(cx)_1 = x_1 ^ (c_1 & x_0)`). Modular addition by the exact carry-mask automaton (per bit a carry-mask bit; checked against brute force at n = 8 on 500 mask triples, max error 0) | exact family |
| Linear, general | The same with the multiply as its shift-and-add decomposition (one adder per set bit of `MUL`, the low known-zero bits of a shifted copy transparent), each adder under the automaton | trail model; over-optimistic for the attacker (section 4.3) |
| Rotational-XOR | Measured on the real code: for every rotation r in 1..31 and k = 1..4, the per-bit bias of `rot_r(M^k(x)) ^ M^k(rot_r(x))` over 2^20 states, the largest |z| of the 512 bits, and the count of exact rotational pairs; plus the word-level prologue `g(x) = (x ^ C) * MUL` alone: the most frequent value of `rot_r(g(x)) ^ g(rot_r(x))` over 2^20 inputs | measurement |
| The fold | The identities a chip would need to pay less than k x 128 for k applications, each tested on 2^20 random inputs, plus the algebraic argument (section 4.5) | measurement and argument |
Search: for each (model, day, k = 1..4) the weight bound W is probed upward (SAT means a trail of weight at most W
exists, UNSAT means none does in the model), then narrowed to the minimum. A k-application trail restricted to one
application is a valid 1-application trail, so every application is held to the proven k = 1 minimum of the same
model (the Matsui floor in the tables). Solver CaDiCaL 1.9.5 through python-sat 1.9. Every trail found of
measurable weight is measured on the real code before it counts: per application and as a chain, 2^20 to 2^28
samples (`attack-f2 verify-diff` / `verify-lin`), with the multiply-layer word transitions counted exactly over all
2^32 inputs (`verify-mults`). A trail that does not hold is blocked and the solver asked again at the same bound.
Linear trails whose correlation cancels inside one adder's hull are caught first by the exact signed sum over the
adder's carry masks.
What "reaches k applications" means here, two readings: (a) the shortcut reading, the one with a cost consequence:
a relation of probability 1 (weight 0) over k applications, which a chip could use to skip work; (b) the
distinguisher reading: a trail of weight under 64 over k applications, the usual practical line. For the gate both
are reported.
## 3. Harness
| Item | Path |
|---|---|
| Crate (ground truth: parameters, vectors, verification, RX, fold) | `tools/attack/f2-mixer/` (`Cargo.toml`, `src/main.rs`), `igneum-pow` by path, own `[workspace]` |
| SAT models and the search | `tools/attack/f2-mixer/model.py` (`selftest`, `search`, `show`) |
| Box queue runner, tables | `tools/attack/f2-mixer/run_jobs.sh`, `tools/attack/f2-mixer/summarise.py` |
| Build line (from the crate directory) | `IGNEUM_AGENT=attack-f2 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f2" --out <scratch> -- build --release`; binary on the box `/srv/builds/igneum-wt-attack/tools/attack/f2-mixer/target/release/attack-f2` (ELF x86-64, sha256 `150337ec...`, the third build; 12 s incremental) |
| Box scratch (states, trails, logs, venv) | `/srv/builds/igneum-wt-attack/target-attack-f2/` (`state/`, `logs/`, `params/`, `vectors/`, `venv/`). The brief's path `attack-f2/` was wiped within ten minutes by another lane's worktree-root rsync (`--delete` spares only `target-*`), so the scratch moved under a `target-` name, as F1 and F6 did |
| Run lines | `venv/bin/python3 model.py selftest --vectors vectors --params-dir params`; `bash run_jobs.sh jobs.txt 11` (each job `model.py search --kind diff|lin --family msb|general|low2 --params params/<day>.<variant>.txt --apps k --state state/<name>.json --budget <chunk> --per-app-min <k=1 floor> --verifier <binary>` under `flock -s /srv/builds/_locks/measure`, `nice -n 10 taskset -c 6-11,54-59`); `attack-f2 rx --day D --variant V --apps 4 --log2 20`; `attack-f2 rx-word --day D --log2 20`; `attack-f2 fold --day D --log2 20` |
| Logs | `logs/<model>-<day>-<variant>-k<k>.log` per search, `logs/rx.<day>.<variant>.log`, `logs/rx-word.<day>.<variant>.log`, `logs/fold.<day>.log`, `logs/summary.md` (the tables below), `logs/verify*.log` |
## 4. Results
### 4.1 The harness fires (known pass, known fail)
| Case | Expected | Got | Log |
|---|---|---|---|
| Selftest: evaluator against `attack-f2 vectors`, 3 days x 3 variants, 4 applications x 16 states each | all match | 9 of 9 files, 64 of 64 applications each | `selftest` output, `logs/selftest.log` |
| Selftest: linear add automaton against brute force, n = 8 | exact | 300 random triples and 200 shifted-copy triples, max error 0.00e+00 | same |
| Selftest: Lipmaa-Moriai against brute force, n = 8; both SAT encodings against their rules at n = 32 | exact | max error 0; 0 mismatches of 40 and 40 | same |
| Selftest: the multiply model's word cost against the sampled best transition (word 3, genesis day) | MSB exact; others within a few bits | MSB: weight 0, measured 2^-0 (exact); bit 30: model 2, sampled best 2^-1.00; bits 31+5: model 8, sampled best 2^-6.03; bit 0: model 10, sampled best 2^-8.97 | same |
| Known pass, 0 applications | the identity trail, weight 0 | trivial (input = output, no weights); not run as a job | |
| Known fail, `rot0` (every rotation 0), differential, k = 1, 2, 4 | a weight-0 trail (MSB-only differences stay MSB-only when nothing rotates) | weight 0 found at k = 1, 2, 4 (both families); measured probability 1 on the real code (`verified_chain -0.0`) | `state/diff-msb-2026-10-03-rot0-k{1,2,4}.json`, `state/diff-general-2026-10-03-rot0-k{1,2}.json` |
| Known fail, `rot0`, linear, k = 1, 2, 4 | a weight-0 trail (LSB masks) | weight 0 at k = 1, 2, 4; measured correlation 1 per application and as a chain | `state/lin-low2-2026-10-03-rot0-k{1,2,4}.json`, `state/lin-general-2026-10-03-rot0-k{1,2}.json` |
| Known fail, `nomul` (MUL 1, RC 0, rk 0: the bare double round), rotational-XOR, k = 1 | a large per-bit bias | max |z| 134.2 (r = 31) against 4.2 for the real mixer; word-level: the prologue is exactly rotational (2^20 of 2^20) against 3 of 2^20 | `logs/rx.2026-10-03.nomul.log`, `logs/rx-word.2026-10-03.nomul.log` |
| Known fail, `rot0`, rotational-XOR | bias | max |z| 32.3 at k = 1, 9.0 at k = 2 | `logs/rx.2026-10-03.rot0.log` |
| Known fail, `nomul`, differential k = 1 | the bare double round's best trail, below the real mixer's | weight 7 found (model), measured 2^-5.0 on the real code | `state/diff-general-2026-10-03-nomul-k1.json` |
### 4.2 Differential trails
| Model | Day | Variant | k | Best trail weight found | No trail at or below (model) | Closed | Per-application floor | Verified on the real code (chain; per application) | Solver s |
|---|---|---|---|---|---|---|---|---|---|
| diff/general | 2026-10-03 | real | 1 | 12 | 11 | yes | 0 | 12.011; [11.939] | 186 |
| diff/general | 2026-10-03 | real | 2 | none | 24 | no (timebox) | 12 | | 1,739 |
| diff/general | 2026-10-04 | real | 1 | 10 | 9 | yes | 0 | 10.001; [9.999] | 321 |
| diff/general | 2026-10-04 | real | 2 | none | 20 | no (timebox) | 10 | | 663 |
| diff/general | 2027-03-01 | real | 1 | 12 | 11 | yes | 0 | 12.057; [11.907] | 175 |
| diff/msb | 2026-10-03 | real | 1 | 12 | 11 | yes | 0 | 12.206; [11.972] | 4 |
| diff/msb | 2026-10-03 | real | 2, 3, 4 | none | 512 (the family dies) | yes | 12 | | 26, 33, 22 |
| diff/msb | 2026-10-04 | real | 1 | 10 | 9 | yes | 0 | 10.001; [10.001] | 309 |
| diff/msb | 2026-10-04 | real | 2, 3, 4 | none | 512 | yes | 10 | | 20, 33, 44 |
| diff/msb | 2027-03-01 | real | 1 | 12 | 11 | yes | 0 | 12.057; [11.907] | 3 |
| diff/msb | 2027-03-01 | real | 2, 3, 4 | none | 512 | yes | 12 | | 10, 15, 21 |
| diff/general | 2026-10-03 | nomul (known fail) | 1 | 7 | 6 | yes | 0 | 5.002; [5.003] | 38 |
| diff/general | 2026-10-03 | nomul | 2 | none | 20 | no | 7 | | 1,309 |
| diff/general, diff/msb | 2026-10-03 | rot0 (known fail) | 1, 2, 4 | 0 | | yes | 0 | probability 1 | under 1 |
The general model's k = 3 and k = 4 jobs (closed 14:3x UTC, every job at its 7,200 s cap, `logs/summary.md`):
| Model | Day | k | Best trail found | No trail at or below (model) | Per-application floor | Solver s |
|---|---|---|---|---|---|---|
| diff/general | 2026-10-03 | 3 | none | 35 | 12 | 7,201 (cap) |
| diff/general | 2026-10-03 | 4 | none | 47 | 12 | 7,359 (cap) |
| diff/general | 2026-10-04 | 3 | none | 29 | 10 | 7,350 (cap) |
| diff/general | 2026-10-04 | 4 | none | 39 | 10 | 7,279 (cap) |
| diff/general | 2027-03-01 | 3 | none | 35 | 12 | 7,321 (cap) |
| diff/general | 2027-03-01 | 4 | none | 47 | 12 | 7,284 (cap) |
| lin/general | 2026-10-03 | 3 | none | 24 | 1 | 7,953 (cap) |
| lin/general | 2026-10-03 | 4 | none | 24 | 1 | 7,352 (cap) |
| lin/general | 2026-10-04 | 3 | none | 28 | 1 | 7,373 (cap) |
| lin/general | 2026-10-04 | 4 | none | 24 | 1 | 7,393 (cap) |
| lin/general | 2027-03-01 | 3 | none | 24 | 1 | 7,402 (cap) |
No trail of weight under 32 at three applications (the finding line): the bound reached is 29 to 35 at three and
39 to 47 at four for differentials, 24 to 28 at three and 24 at four for linear masks, all solver-capped, so these are
effort bounds, not proofs; they grow with k as the per-application floors predict.
### 4.3 Linear trails
| Model | Day | Variant | k | Best trail weight found (correlation 2^-w) | No trail at or below | Closed | Verified (chain; per application) | Solver s |
|---|---|---|---|---|---|---|---|---|
| lin/general | 2026-10-03 | real | 1 | 1 | 0 | yes | 0.996; [0.997] | 5 |
| lin/general | 2026-10-03 | real | 2 | none | 20 | no (timebox) | | 1,019 |
| lin/general | 2026-10-04 | real | 1 | 1 | 0 | yes | 0.995; [0.999] | 5 |
| lin/general | 2026-10-04 | real | 2 | none | 24 | no (timebox) | | 1,669 |
| lin/general | 2027-03-01 | real | 1 | 1 | 0 | yes | 1.003; [0.996] | 5 |
| lin/general | 2027-03-01 | real | 2 | none | 20 | no (timebox) | | 1,224 |
| lin/low2 | 2026-10-03, 2026-10-04, 2027-03-01 | real | 1 | 1 | 0 | yes | 0.995 to 1.003 | 4 to 5 |
| lin/low2 | 2027-03-01 | real | 2, 3, 4 | none | 512 (the family dies) | yes | | 12, 18, 27 |
| lin/general | 2026-10-03 | nomul (known fail) | 1 | 1 | 0 | yes | 0.999; [1.003] | 4 |
| lin/general, lin/low2 | 2026-10-03 | rot0 (known fail) | 1, 2, 4 | 0 | | yes | correlation 1 | 5 to 10 |
One application carries a weight-1 linear trail (the LSB mask through the prologue and one add, correlation 1/2),
the structural residue of 4.5; at two applications no trail at or below weight 20 to 24 exists in the general
model within the timebox, and the LSB family dies (no trail at or below 512) from k = 2.
### 4.4 Rotational-XOR
Per k and day, the largest |z| over all 31 rotations and 512 bits at 2^20 states (15,872 bit tests per k; the
noise ceiling of that many tests is about 4.3), and the count of exact rotational pairs.
| Day | k = 1 | k = 2 | k = 3 | k = 4 | Exact pairs | Log |
|---|---|---|---|---|---|---|
| 2026-10-03 | 4.22 (r 19) | 4.29 (r 30) | 4.22 (r 3) | 4.62 (r 27) | 0 | `logs/rx.2026-10-03.real.log` |
| 2026-10-04 | 4.35 (r 17) | 4.00 (r 19) | 4.24 (r 25) | 4.49 (r 28) | 0 | `logs/rx.2026-10-04.real.log` |
| 2027-03-01 | 3.96 (r 7) | 4.07 (r 2) | 4.04 (r 28) | 4.17 (r 19) | 0 | `logs/rx.2027-03-01.real.log` |
| 2026-10-03, bare double round (`nomul`) | 134.24 (r 31) | 4.37 | | | 0 | `logs/rx.2026-10-03.nomul.log` |
The word-level prologue `(x ^ C) * MUL`: over 2^20 inputs the most frequent value of `rot_r(g(x)) ^ g(rot_r(x))`
occurs at most 3 times for every word and every r on all three days (`logs/rx-word.<day>.real.log`, the
`rxw_worst` lines), against 2^20 of 2^20 without the multiply. The odd multiply by a random constant is not
rotational to any measurable degree, and one application already shows no per-bit bias. Rotational-XOR does not
reach 1 application.
### 4.5 The fold of the multiply layer
One application is `D o P_rk`, with `P_rk(s)_i = (s_i ^ (RC_i + rk)) * MUL_i` and `D` the double round (fixed per
day). Multiplication by an odd constant distributes over modular addition and over nothing else in `D` (XOR,
rotation); the XOR with a constant commutes with XOR and rotation and with nothing else (addition, multiply). A fold
across applications would need one of the identities below. Each was tested on 2^20 random inputs on every day
(`logs/fold.<day>.log`):
| Identity a chip would need | Holds on | Meaning |
|---|---|---|
| `(xa ^ Ca) * ma + (xb ^ Cb) * mb = ((xa ^ Ca) + (xb ^ Cb)) * ma` for the four column pairs (0,4), (1,5), (2,6), (3,7) | 0 of 1,048,576 for every pair on every day (`MUL` distinct in every pair) | the multiply does not fold into the first add of a quarter round; it would if a column pair drew the same `MUL` (probability 2^-31 per pair per day, the weak-day class of F4) |
| `(x ^ C) * m = (x * m) ^ (C * m)`, or `= (x * m) ^ C'` for any single `C'` | 0 of 1,048,576; the best single `C'` agrees on 33 of 1,048,576 (2^-15) | the constant cannot be moved past the multiply, so application j + 1's prologue cannot share application j's multiply |
| an XOR constant on one word commuting with the bare double round (so the next prologue's constant could be folded back) | 0 of 65,536 for every word | every word's value feeds an add inside the double round |
| the MSB passing the prologue and the add for free; the LSB passing the prologue | 1,048,576 of 1,048,576 each | the structural residue: the only free passages, both moved by the rotations (the family deaths in 4.2 and 4.3) |
So k applications cost k times one application, 128 hoisted operations each (16 multiplies, 32 adds, 32 XORs, 32
rotations with the constants hoisted); `chip-model-v3.md` 5.2's 9,360 per item stands. The trail weights of 4.2
and 4.3 growing with k is the quantitative side of the same fact: a composition that collapsed to one application's
shape would keep one application's trail weights.
## 5. Gate and verdict
Gate (plan 4.2 F2, 1.4 (1)): no distinguisher or shortcut beyond 2 of the 8 applications between dependent reads,
after the stated search.
| Line of attack | Reach | Verdict |
|---|---|---|
| Differential, general model (Markov on the multiply, exact add rule, SAT) | one application: best trail weight 10 to 12 on three days, verified on the real code; two applications: no trail at or below weight 20 to 24 within 7,200 s per job (not closed); the MSB family dies at two applications on every day | nothing reaches 2 applications below 2^-20 |
| Linear, general model (piling-up, SAT) | one application: weight 1 (the LSB residue); two applications: no trail at or below 20 to 24 within the timebox; the LSB family dies at two | nothing reaches 2 applications below 2^-20 |
| Rotational-XOR | no per-bit bias at one application (max abs z 4.0 to 4.6 at 2^20 states, noise ceiling 4.3); 0 exact pairs; the multiply prologue is rotational on at most 3 of 2^20 inputs; the bare double round fires at 134 | does not reach 1 application |
| Algebraic fold of the multiply layer | every identity a fold needs holds on 0 of 2^20 inputs on every day; k applications cost k | no shortcut |
Verdict: PASS with the effort bound stated: about 60 solver jobs, 2 to 29 minutes each, on three day keys; the
reduced-round margin reached is one application fully characterised (weights 10 to 12 differential, 1 linear) and
two applications with no trail under weight 20 to 24, three with none under 29 to 35 (differential) and 24 to 28
(linear), four with none under 39 to 47 and 24, against 8 applications between reads, so the margin between what the
search reaches and what the construction uses is at least 4 applications at the solver's cap. What this does not do is in section 7; the lower bound is the paid question.
## 6. Consequences per tier
No shortcut, so no tier moves: a home card, a rig and a pool pay the 72 applications per item the verifier pays;
a chip with a fixed datapath pays them too (the fold test), which is what `chip-model-v3.md` 5.2's 9,360 ops per
item assumes. `mixer_mult` stays 8; the verifier measurement of F6 stands unchanged.
## 7. What this does not do
- It does not bound the mixer from below: the general models are trail models (Markov for the multiply's
differential, piling-up for the linear), and the family models are exact only inside their families. The adversarial
lanes' job (funding.md B5 rank 1) is the effort-bounded version of the same search with their own tools.
- Three days, not a census: the ROT, MUL, RC classes over 2^24 days are F4's row. One cheap addition for F4 from
this harness: the MSB-family death at k = 2 (`model.py search --kind diff --family msb --apps 2`) runs in seconds
per day, and a day where it does not die is a weak day of the kind the gate is about.
- Differential and linear only, as the row says: no boomerang, no integral or cube property, no related-key (the
round keys are public constants).

View file

@ -0,0 +1,148 @@
# F3: the chained cache's j + 1 bound and the storage-against-recompute curve
Attack-pass row F3 of `docs/plans/cryptanalysis.md` section 4.2 (the record is `docs/analysis/attack-pass-2026-10.md`). Run 7 October 2026, 09:10 to 09:12 UK (08:10 to 08:12 UTC in the logs), on igneum-build-1. Verdict: PASS on all three gate clauses. No line (s, j) is derivable in fewer than j + 1 block evaluations without an earlier line, by an exhaustive search over the block dependency graph extracted from the code at 64 and 1,024 lines, cross-checked by an exhaustive pebbling search over every configuration at 10 lines. The storage-against-recompute curve over cache lines is monotone from f = 1/64 to 1. The f = 1 point is unchanged.
## Target
| Item | Value |
|---|---|
| Commit | 924288d1 (the brief); the worktree HEAD moved to 11b375a0 during the run; `igneum-pow/src/memhard.rs` is byte-identical at both (blob ad42470b, `git diff --stat 924288d1 HEAD -- igneum-pow/src/memhard.rs` empty) |
| Construction, from the code | `Cache::fill_segment`: `in_j = prev XOR (sigma || K || seg || j || tag)`, `line_j = chacha_block(in_j)` where `chacha_block(x) = ChaCha12core(x) + x`, `prev_0 = 0`, `prev_j = line_{j-1}`; 64 lines per segment, 2^16 segments, 2^26 words (256 MiB) |
| Reads | `derive_items_mask`: 8 dependent reads per item at line index `s[0] AND mask`, so the segment and j of a read are uniform over the 2^22 lines (F8 checks the uniformity) |
| Known-failed shape | a line (s, j) computable in fewer than j + 1 block evaluations without an earlier line of segment s (the address-steering shape of the MTP break, Dinur and Nadler 2017, needs a data-dependent chain; this chain's inputs are fixed by the key, so the shape to search is a structural shortcut on the dependency graph) |
| Gate | no derivation under j + 1 blocks; the curve monotone; the f = 1 point unchanged |
| Prior evidence | none to re-gate: the `ca2-cache` branch named in the status board is the hot-table experiment (`docs/plans/hot-table.md`), not a chain analysis |
## Method
The model is the code, not the prose. `tools/attack/f3-cache/src/main.rs` runs one chain function, written in the shape of `memhard.rs` (the quarter round, the 6 double rounds, the feed-forward, the prev XOR, the constant block), generically over two word types:
| Word type | What it computes | Use |
|---|---|---|
| `u32` | the real arithmetic | `verify`: bit-exact against `Cache::fill_segment` on 16 (key, segment) pairs and against `chacha_block` on 100,000 random inputs |
| taint set | which block outputs a value depends on (add, xor, rotate = union) | `search`: the direct-parent graph of every block, with each computed line relabelled to the single node {j} so parents are direct, not transitive; plus the 16 x 16 (output word, input word) dependency matrix of one block |
The exhaustive search: for every target line j, the minimum number of block evaluations with nothing stored is the size of the backward closure of j on the extracted graph (every non-stored block in the closure must be evaluated at least once; once each in dependency order suffices). `pebble` checks that formula against an exhaustive 0-1 BFS over every pebble configuration (place on a node whose parents are pebbled at cost 1, remove at cost 0) for all 2^10 stored sets x 10 targets on each of the three graphs: 10,240 pairs per graph, 0 mismatches. Two deliberately broken chains are the known-fail cases: `skip2` (line j fed from line j - 2) and `nofeed` (no previous line fed in). The curve: for f = 1/64 to 1 (fraction of cache LINES held), the blocks per read on the naive pattern of `funding.md` B2 rank 2 (every L/n-th line from line 0) and on the optimal pattern (exact DP over chunk lengths; brute force over every C(64, n) set for n up to 8, 4,426,165,368 sets at n = 8); ops per item = 9,360 mixer ops (chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block (608 counted from the code: 48 quarter rounds x 12, 16 feed-forward adds, 16 input XORs; also at MEMHARD.md's approximate 700). Three more checks on the real function: single-bit avalanche and a differential-independence test on `chacha_block`, a census of every line of the real 2^22-line cache for the day key 2026-10-03, and one-core timings of a block, a mixer application, an item and a line recompute.
## Harness
| Item | Value |
|---|---|
| Crate | `/Users/joshm/Projects/igneum-wt-attack/tools/attack/f3-cache/` (`Cargo.toml` with `igneum-pow = { path = "../../../igneum-pow" }` and an empty `[workspace]`; `src/main.rs`; `run-box.sh`) |
| Build | `cd tools/attack/f3-cache && IGNEUM_AGENT=attack-f3 IGNEUM_TOOLCHAIN_MISMATCH=ok bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f3" --out <scratchpad>/attack-f3 -- build --release` (rc 0, 38 s wall, 0 warnings; the Mac's PATH rustc is 1.69 but `~/.cargo/bin/rustc` is 1.99.0, which the script read as "on both sides"; log `<scratchpad>/attack-f3/build-1.log`) |
| Binary | box `/srv/builds/igneum-wt-attack/tools/attack/f3-cache/target/release/attack-f3`, sha256 975115a385ca3195...71cc33, 563,104 bytes |
| Run | on the box: `nohup bash run-box.sh r1 > run-r1.log 2>&1 &` from `/srv/builds/igneum-wt-attack/attack-f3/`; every phase as `flock -s /srv/builds/_locks/measure -c "nice -n 10 taskset -c 12-15,60-63 attack-f3 <cmd>"`, one chunk per phase, the whole run 17 s (08:10:58 to 08:11:15 UTC; box load 54 at start) |
| Phase lines | `verify`; `search --lines 64|1024 --variant real|skip2|nofeed`; `pebble --lines 10`; `store --lines 64 --brute-max 8`; `store --lines 1024 --brute-max 2`; `curve --lines 64`; `curve --lines 64 --ops-block 700`; `curve --lines 1024`; `avalanche --samples 1048576`; `census --day 2026-10-03`; `bench --n 20000000` |
| Logs | box `/srv/builds/igneum-wt-attack/attack-f3/run-r1.log` and `r1-<phase>.log`; Mac copies `/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/attack-f3/` |
Box hygiene: the box checkout of every `build-remote.sh` run on this worktree executes `git clean -fd` at `/srv/builds/igneum-wt-attack` (remote-run.sh `checkout_tree`), which deletes any untracked scratch directory there. `attack-f3/` and `attack-f3-venv/` are listed in that mirror's `.git/info/exclude` so they survive; nothing in the tree was touched. The F1 lane's `attack-f1-venv/` is untracked and unprotected and will be removed by the next build from any agent on this worktree.
## The two firings and the pass
| Chain | Direct parents (taint trace) | Lines under j + 1 at 64 lines | Cheapest derivations | Exhaustive pebbling at 10 lines, cost per target | Verdict | Log |
|---|---|---|---|---|---|---|
| real (the code) | j - 1 for all 63 lines after line 0 | 0 of 64 | none; every line costs exactly j + 1 (mean 32.5) | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 | PASS | `r1-search-real-64.log`, `r1-pebble-10.log` |
| real, 1,024-line model | j - 1 for all 1,023 lines after line 0 | 0 of 1,024 | none (mean 512.5) | same graph rule | PASS | `r1-search-real-1024.log` |
| skip2 (known fail A) | j - 2 for 62 lines, none for 2 | 63 of 64 | j = 1 in 1, j = 63 in 32 (mean 16.5) | 1, 1, 2, 2, 3, 3, 4, 4, 5, 5 | FIRE | `r1-search-skip2-64.log`, `r1-search-skip2-1024.log` |
| nofeed (known fail B) | none for all 64 | 63 of 64 | every line in 1 block (mean 1.0) | 1 x 10 | FIRE | `r1-search-nofeed-64.log`, `r1-search-nofeed-1024.log` |
`verify` (`r1-verify.log`): the model chain equals `Cache::fill_segment` on keys {day 2026-10-03, 3 random} x segments {0, 1, 12345, 65535} (16 of 16), `block == chacha_block` on 100,000 of 100,000 random inputs, the 1,024-line model's first 64 lines equal the 64-line chain, and both broken variants differ from the real chain from line 1 (line 0 equal, as the rule predicts). The block's word dependency matrix is full on every variant (256 of 256 pairs), so the firings come from the chain rule alone.
## Derivation cost per line on the model segment (real chain, nothing stored)
| j | blocks to derive line j | j + 1 | Log |
|---|---|---|---|
| 0 | 1 | 1 | `r1-search-real-1024.log` |
| 1 | 2 | 2 | |
| 3 | 4 | 4 | |
| 7 | 8 | 8 | |
| 15 | 16 | 16 | |
| 31 | 32 | 32 | |
| 63 | 64 | 64 | (the last line of a real segment; `r1-search-real-64.log` lists all 64) |
| 127 | 128 | 128 | |
| 255 | 256 | 256 | |
| 511 | 512 | 512 | |
| 1,023 | 1,024 | 1,024 | |
All 1,024 lines were searched (0 under j + 1, mean 512.5 = (L + 1) / 2); the 64-line table in `r1-search-real-64.log` has every j from 0 to 63 at exactly j + 1.
## Store patterns on the real 64-line segment
Blocks per read averaged over j uniform in 0..63. "Naive" is `funding.md` B2 rank 2's pattern (every k-th line from line 0). "Optimal" is the exact minimum over store sets of that size (DP; brute force over every set for n up to 8, agreeing with the DP on every row it ran). The gap formula equals the closure cost on the extracted graph on 2,000 of 2,000 random stored sets (`r1-store-64.log`).
| f | Stored lines n | SRAM held | Naive blocks per read | Optimal positions | Optimal blocks per read | Brute force over C(64, n) sets |
|---|---|---|---|---|---|---|
| 1/64 | 1 | 4 MiB | 31.5 | [32] | 16.0 | 16.0 (64 sets) |
| 1/32 | 2 | 8 MiB | 15.5 | [21, 43] | 10.5 | 10.5 (2,016 sets) |
| 1/16 | 4 | 16 MiB | 7.5 | [12, 25, 38, 51] | 6.094 | 6.094 (635,376 sets) |
| 1/8 | 8 | 32 MiB | 3.5 | [7, 15, 22, 29, 36, 43, 50, 57] | 3.172 | 3.172 (4,426,165,368 sets, 11.2 s) |
| 1/4 | 16 | 64 MiB | 1.5 | [3, 7, 11, ..., 55, 58, 61] | 1.453 | not run (DP exact) |
| 1/2 | 32 | 128 MiB | 0.5 | odd lines | 0.5 | not run |
| 1 | 64 | 256 MiB | 0 | all | 0 | not run |
The 1,024-line model (`r1-store-1024.log`) gives 29.68 / 15.05 / 7.40 / 3.48 / 1.50 / 0.5 / 0 at the same f on the optimal pattern: the naive and optimal patterns converge as the chain lengthens, because the wasted stored line 0 and the end effects are a smaller share.
## The curve: ops per item against the fraction of cache lines held (real 64-line segment)
Ops per item = 9,360 (the 72 mixer applications, hoisted, plus the fold: chip-model-v3.md 5.2) + 8 reads x blocks per read x ops per block. `r1-curve-64.log` (608 ops per block, counted) and `r1-curve-64-memhard.log` (700, MEMHARD.md item 4). Ops per hash = 128 x ops per item + 512. MH/s at the chip model's 50 T op/s budget (approximate, chip-model-v3.md section 1).
| f (lines held) | SRAM | Blocks per read, naive / optimal | Ops per item, naive, 608 | Ops per item, optimal, 608 | Ops per item, optimal, 700 | Ops per hash, optimal, 608 | MH/s at 50 T op/s, optimal, 608 |
|---|---|---|---|---|---|---|---|
| 1/64 | 4 MiB | 31.5 / 16.0 | 162,576 | 87,184 | 98,960 | 11,160,064 | 4.5 |
| 1/32 | 8 MiB | 15.5 / 10.5 | 84,752 | 60,432 | 68,160 | 7,735,808 | 6.5 |
| 1/16 | 16 MiB | 7.5 / 6.094 | 45,840 | 39,000 | 43,485 | 4,992,512 | 10.0 |
| 1/8 | 32 MiB | 3.5 / 3.172 | 26,384 | 24,788 | 27,122 | 3,173,376 | 15.8 |
| 1/4 | 64 MiB | 1.5 / 1.453 | 16,656 | 16,428 | 17,498 | 2,103,296 | 23.8 |
| 1/2 | 128 MiB | 0.5 / 0.5 | 11,792 | 11,792 | 12,160 | 1,509,888 | 33.1 |
| 1 | 256 MiB | 0 / 0 | 9,360 | 9,360 | 9,360 | 1,198,592 | 41.7 |
Monotone: ops per item is non-increasing in f on both patterns at both op counts and on the 1,024-line model (`CURVE ... monotone non-increasing` in all three curve logs). The f = 1 point: 9,360 ops per item, 1,198,592 ops per hash, 41.7 MH/s at 50 T op/s, which is the chip-model-v3.md section 5.4 row "none, f = 0" of the published ITEM curve (the on-die-cache recompute chip of sections 1 to 3). The published item curve stores dataset ITEMS and is a different curve: its f = 1 point (GDDR7, 166.4 MH/s, 0.466 microjoules per hash) contains no cache read and no mixer op, so nothing in this row touches it. `funding.md` B2 rank 2's arithmetic reproduces on the naive pattern at 700 ops per block: 3.5 blocks per read, 2,450 ops per line, 19,600 per item on top of the mixer, 13.5 MH/s (50 T / (128 x 28,960 + 512)).
## Measured times, one box core (`r1-bench.log`, `r1-census.log`; nice 10, cores 12-15,60-63, box load 54)
| What | Measured | Note |
|---|---|---|
| One ChaCha12 block, dependent chain of 20,000,000 | 66.64 ns | |
| One mixer application (class v4 parameters), dependent chain of 20,000,000 | 17.29 ns | block / application = 3.85 (counted ops 608 / 128 = 4.75) |
| One item against the 256 MiB cache, batches of 32 | 1,326 ns | 72 applications = 1,245 ns; the 8 dependent reads and the fold add 81 ns because the batch overlaps them |
| One line recomputed from nothing, 312,500 random (seg, j) | 2,734 ns | 32.5 blocks per line on average, 84.1 ns per block inside the chain |
| The 256 MiB cache fill, one thread | 0.36 to 0.4 s | 86 ns per block with the writes |
In measured time, holding every 8th line at the optimal placement makes an item cost 72 + 8 x 3.172 x 3.85 = 170 mixer-application equivalents against 72, a 2.36x penalty per item (2.65x in counted ops). Holding one line in 64 costs 72 + 8 x 16 x 3.85 = 565, a 7.8x penalty.
## Checks on the real function (`r1-avalanche.log`, `r1-census.log`)
| Check | Result |
|---|---|
| Single-bit avalanche of `chacha_block`, 1,048,576 flips | mean 256.00 of 512 output bits change (ideal 256), min 200, max 312 |
| (output word, input word) pairs where an output word did not change | worst count 0 of 1,048,576 |
| Chain step: one bit of line j - 1 flipped | line j changes 255.96 bits, line j + 1 changes 255.94 (131,072 flips) |
| Differential independence: B(x ^ d) ^ B(x) == B(y ^ d) ^ B(y) over 262,144 (x, y, single-bit d) | 0 cases |
| Census of the real cache, day key 2026-10-03 | 4,194,304 of 4,194,304 lines distinct, 0 all-zero lines: no two chains merge and no block input repeats |
## Gate
| Clause | Result | Where |
|---|---|---|
| No derivation under j + 1 blocks | 0 of 64 and 0 of 1,024 lines under j + 1 on the extracted graph; the formula exact on 10,240 of 10,240 exhaustive pebbling cases; both known-fail chains fire | `r1-search-real-64.log`, `r1-search-real-1024.log`, `r1-pebble-10.log` |
| The curve monotone | non-increasing on both patterns, both op counts, both segment lengths | the three `r1-curve-*.log` |
| The f = 1 point unchanged | 9,360 ops per item = chip-model-v3.md 5.4 "none, f = 0" row; the item curve's GDDR7 f = 1 row (166.4 MH/s, 0.466 microjoules) untouched | `r1-curve-64.log` |
Verdict: PASS.
## Observations that are not findings
| Observation | Number | What it means | What I propose |
|---|---|---|---|
| `funding.md` B2 rank 2 prices the honest trade-off at the naive placement | 3.5 blocks per read at f = 1/8 against 3.17 optimal (9.4 percent less); 31.5 against 16.0 at f = 1/64 (2.0x less, because storing line 0 is worthless: it costs 1 block anyway) | the chip at f = 1/8 reads 15.8 MH/s (608 ops per block, optimal placement) or 14.4 (700, optimal) against `funding.md`'s 13.5 (700, naive); still 0.38x of the full SRAM mirror's 41.7 and 0.12x of the 5090's 136.1 (chip-model-v3.md section 2); the curve stays monotone, so the published verdict (the partial chip is not the threat, the full mirror beats it) stands | one sentence in `funding.md` B2 rank 2: "holding every 8th line at the best placement costs 3.2 blocks per read (3.5 for every 8th line from line 0)". Not edited here: outside this row's two files; for main to serialise |
| The chain's hardness per line is sequential time, not memory | one pebble (64 bytes) over j + 1 steps: the cumulative memory of deriving a line is about 64 x (j + 1) byte-steps | the chain protects the cache by op count, which is exactly what the curve prices in ops; parallel attackers pipeline items and pay E(f) x 608 ops per read in throughput, E(f) block latencies in latency; a chip that holds nothing (f = 0) pays 32.5 x 608 = 19,760 ops per read, 158,080 per item, 167,440 with the mixer (17.9x the mixer alone), 2.3 MH/s at 50 T op/s | nothing to move; the public model should keep quoting ops, never bytes, for this piece |
| What this row does not cover | a cryptanalytic shortcut inside `chacha_block` in this chaining mode (the differential and avalanche tests are sanity checks, not a bound) | the adversarial lanes's rank 2 question (`funding.md` B2) stays worth the money; plan 4.2 says the internal pass cannot prove the chain's trade-off curve | none |
## Consequences per user tier
| Tier | What this row changes |
|---|---|
| Home miner, one 8 / 12 / 16 / 24 or 32 GB card, any vendor, any OS | nothing: the honest miner holds the dataset, the verifier holds the 256 MiB cache; no memory, hash rate, or power figure moves |
| Rig, pool user | nothing |
| Chip builder | the partial-cache chip is priced 9 percent better at f = 1/8 and 2x better at f = 1/64 than `funding.md` says, and is still worse than the full SRAM mirror at every f below 1; the public per-joule sentence (evidence row 17, 2.1x at k = 1) rests on the item curve's f = 1 point, which this row leaves untouched |
| The public report | the record and the harness are published; rank 2's open question is the block function in chaining mode, not the graph |

View file

@ -0,0 +1,339 @@
# F4. The weak-day census: 2^24 day keys through `MixParams::with_shape`
Attack pass row F4 (`docs/plans/cryptanalysis.md` section 4.2; the gate is section 1.4 (3) and `funding.md` B5
rank 3; the threat is `funding.md` B2 rank 3). Run 7 October 2026, 09:10 to 09:55 UK, on igneum-build-1 by the
attack-f4 agent (the verifier timing row of 6.6 queued behind other lanes' holds). Every number below cites its log.
## Verdict
**PASS on the gate read against M2, the DSP-bound per-day datapath (0 days over 1.1x in 2^28), and on every named
weak class; the generous bound M1 (every multiply in LUT adders) exceeds the gate at 3.26e-4 of days as the tail of a
sum, not a class, and is routed to main as a bound finding with a rejection-and-redraw rule for the next class.
Class v4 is not changed.**
Which metric the 1.1x gate reads against, and why: M2. The gate (plan 1.4 (3)) asks for the fraction of days in a
weak class, and M1's excess has no class behind it (section 6.2: the exact 16-fold convolution of one random NAF
weight predicts the census to 0.6 percent). A per-day FPGA attacker who builds the 16 multiplies in LUT shift-add
trees is building the slower design: those trees are 72 percent of M1's cost (167 of 231 adders), and DSP blocks
take that cost off the fabric, so the design that wins is DSP-bound, where the day's constants move nothing unless a
word has NAF weight at most 3, which happens on no day in 2^28 for two words. M1 is still reported in full because
the brief asks for the generous bound, and because a two-line rule closes it for nothing.
| Metric | Days over 1.1x in 2^24 | Fraction | Days over 1.1x in 2^28 | Fraction | Gate 2^-20 = 9.54e-7 | Log |
|---|---|---|---|---|---|---|
| M1: per-day LUT datapath, adders per mixer application, against the census median | 5,476 | 3.264e-4 | 87,426 | 3.257e-4 | OVER, by 342x | `census-2p24.md`, `census-2p28.md` gate table |
| M1 exact expectation (16-fold convolution of the NAF-weight table over all 2^31 odd constants) | 5,441 | 3.243e-4 | | | the census is the tail of a smooth sum, not a class | `expect-231.log` last line |
| M2: DSP-bound datapath, 16/(16 - k), k = words of NAF weight at most 3 | 0 | 0 | 0 | 0 | under | `census-2p24.md`, `census-2p28.md` M2 table |
| ROT value and RC value on a per-day datapath | 0 | 0 | 0 | 0 | under (exact 0 ops moved, section 3) | section 3 |
The gate as written fails under M1 only. What M1 finds is not a weak class: the per-day cost of the 16 constant
multipliers is a sum of 16 NAF weights (mean 231.1 adder-equivalents per application, sd 6.19), and 1 day in 3,070
sits 3.4 sigma below the median, where a bitstream synthesised for that day pays 10 to 19 percent fewer adders. The
worst day in 2^28 reads 1.19x (day 27,952,752, cost 194). The exact expectation predicts the census to 0.6 percent.
Section 7 prices the consequence (0.004 percent more hashes a year for an all-LUT FPGA that re-synthesises every
day, nothing for a chip or a GPU) and section 8 gives the rejection-and-redraw rule that closes it.
## 1. Target
| Item | Value |
|---|---|
| Commit | 924288d1 (the brief); the worktree HEAD moved to 11b375a0 during the pass (F5 and F6 records); `git diff 924288d1 11b375a0 --stat -- igneum-pow/src` is empty, so the target code is the same |
| Code | `igneum-pow/src/memhard.rs` `MixParams::with_shape` (lines 189 to 215): `SplitMix64::new(key[0] as u64 \| (key[1] as u64) << 32)`, then `ROT[0..7] = 1 + below(31)`, `MUL[0..15] = next() as u32 \| 1`, `RC[0..15] = next() as u32`; no rejection rule |
| Day key | `bind::day_bytes(d) = "igneum-day/" \|\| d_le64`, `key = seed_words_from_bytes(day_bytes)` (the interim day rule, `bind.rs` lines 30 to 68); the genesis day index is 20,729 (`bind.rs` test `day_bytes_layout`) |
| Shape | `Shape::for_class(&V4_CLASS)`: mixer x8, cache 2^26 words, no derivation program (asserted by the harness) |
| Mixer | `memhard::mixer`: per word `(s ^ (RC + rk)) * MUL`, then one ChaCha double round with `ROT[0..3]` on the columns and `ROT[4..7]` on the diagonals; 72 applications per item under x8 |
| Census set | 2^24 consecutive chain days from 20,729 (the gate run), and 2^28 (the extended run); the first 36,525 of them are the chain's public calendar for the next 100 years under the interim rule |
The 64-bit seeding fact (F7 covers the spec's intent): the 40 draws depend on `key[0] | key[1] << 32` alone, so the
stream can produce at most 2^64 distinct parameter sets whatever the key's other 192 bits hold. Over the 2^24 census
days the 64-bit seeds were all distinct (0 collisions, expected 7.6e-6; `census-2p24.md` "64-bit seeding" line).
`below(31)` is `next() % 31` without rejection: the bias per rotation value is 2^-64 and is ignored.
## 2. Known-failed shape
A day key whose drawn `ROT`, `MUL` or `RC` gives a fixed datapath a gain over 1.1x: all-equal `ROT` (31^-7 per day,
MEMHARD.md section 3 item 3, untested until now), `MUL = 1` (2^-31 per word), pairs summing to 32, small rotation
amounts, low-weight multipliers, `RC + rk = 0`.
## 3. The gain metrics (exact, structural)
The verifier and every GPU run the same instructions on every day (`rotate_left` by a register amount, `wrapping_mul`,
no branch on a drawn value), so wall time cannot move with the draw; the only attacker a weak day helps is one who
builds the day's constants into logic. That is an FPGA bitstream synthesised per day (hours of compile against a
public calendar), never a taped-out chip. Costs are in 32-bit adder-equivalents per mixer application:
| Element of one application | Generic datapath | Per-day datapath |
|---|---|---|
| 16 x `s ^ (RC + rk)` | 16 | 0 (constant XOR: inverters, absorbed into the next LUT) |
| 16 x `* MUL` | 16 multipliers (value-independent) | M1: `NAF(MUL_i) - 1` adders each (canonical signed-digit shift-add); M2: a DSP block each, value-independent, except a word of NAF weight at most 3 moves to 2 LUT adders and frees its DSP |
| 8 quarter rounds: 32 adds, 32 XORs | 64 | 64 |
| 32 rotations | 32 barrel shifters | 0 (wiring) |
* **M1** `cost = 64 + sum_i (NAF(MUL_i) - 1)`; gain of a day = census median cost / the day's cost. The generous
bound: optimal single-constant multiplication is below NAF for every constant and the ratio between days is what
is measured.
* **M2** gain = `16 / (16 - k)` on a DSP-bound design, k the words of NAF weight at most 3.
* **ROT** and **RC** hand a per-day datapath exactly 0 ops at any value (wiring and inverters); on a generic
datapath a rotation costs the same at every amount and `RC + rk = 0` removes one XOR of 10,368 ops per item
(1.0001x). They are censused as structure, and the worst members are measured for diffusion (section 6), the
only other thing a rotation draw could move; a bit-exact verifier never lets a chip skip an application, so
diffusion is reported and is not a gain.
## 4. Harness
| Item | Path or line |
|---|---|
| Crate | `tools/attack/f4-weakday/` (`Cargo.toml` with `igneum-pow = { path = "../../../igneum-pow" }` and an empty `[workspace]`; `src/main.rs`); `igneum-pow` untouched |
| Build | `cd tools/attack/f4-weakday && IGNEUM_AGENT=attack-f4 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f4" --out <scratch> -- build --release`; box binary `/srv/builds/igneum-wt-attack/tools/attack/f4-weakday/target/release/attack-f4`: sha256 `fda006d7...835f52` ran every census and firing (`build-1.log`); the rebuild `5eb081cf...7f0355` (`build-2.log`) removes one unused import and nothing else |
| Unit tests | `build-remote.sh --no-fetch -- test --release` on the box (`test-1.log`): 2 passed, 0 failed (`naf_weights`: 0, 1, 3, 7, 2^32 - 1, the alternating maximum 17, and the planted weight-3 constant; `genesis_day_draw_matches_memhard_md`: the string day `2026-10-03` draws `ROT 20 20 19 4 26 3 3 27`, MEMHARD.md section 1.1, through the same `with_shape` path the census uses) |
| Census (gate) | `flock -s /srv/builds/_locks/measure -c 'nice -n 10 taskset -c 16-21,64-69 attack-f4 census --from 20729 --count 16777216 --threads 12 --dedupe --out census-2p24.md'`; 4.2 s |
| Census (extended) | the same with `--count 268435456 --out census-2p28.md`; 68.6 s |
| Expectation tables | `attack-f4 expect --threads 12 --median 231` (every odd 32-bit constant: NAF weight and popcount, then the 16-fold convolution); 14.9 s |
| One day | `attack-f4 day --index <d> --median 231` |
| Firings | `attack-f4 plant alleq\|mul1\|mul1all\|mulnaf\|rc0\|rcrk0 --median 231` (the day 20,729 draw with one field forced through the crate's own hook) |
| Diffusion | `attack-f4 avalanche --index <d> --states 2048 [--plant-alleq r]` |
| Timing (exclusive hold) | `timing.sh` on the box under nohup: `flock -x -w 7200 /srv/builds/_locks/measure -c 'nice -n 19 taskset -c 16,64 igneum-pow bench --seed x --epoch-hex edc4fa84...fb07 --day-hex <day bytes> --program-class v4 --warps 100'` for day 20,729 and the worst day, A B A B. Process note: withdrawing the first attempt, one ad hoc ssh line used `pkill -f "<literal>"`, the banned shape, and killed its own shell (self-match); the relaunch used the bracket form. Nothing else was touched |
| Calendar | `attack-f4 census --from 20729 --count 36525 --threads 12 --out census-100y.md` (the chain's first 100 years) |
| Box logs | `/srv/builds/igneum-wt-attack/attack-f4/{run2.log, census-2p24.md, census-2p28.md, census-100y.md, expect-231.log, firings.log, avalanche.log, timing.log}` |
| Mac copies | `/private/tmp/claude-501/-Users-joshm/cd75457f-4858-4f86-9634-7481ee056b7b/scratchpad/attack-f4/box/` (the box directory was deleted once from under the pass at about 09:14 UK by another agent's worktree sync; everything was re-run and copied to the Mac the moment it ended; the re-run reproduced the first run line for line) |
## 5. The two firings (`firings.log`)
| Case | Classifier | Gain | Result |
|---|---|---|---|
| Known-pass: day 20,729 (the genesis day), `ROT [6, 25, 5, 25, 29, 11, 9, 21]`, NAF sum 178 | no weak class (only "pair sums to 32", 12 and 60 percent of all days) | M1 0.978x, M2 1.000x | passes, as it must |
| Known-fail: `plant mul1all` (all 16 `MUL = 1`) | `MUL any = 1` FIRED | M1 3.453x, M2 unbounded | FIRED over 1.1x |
| Known-fail: `plant mulnaf` (four words at NAF weight 3) | `MUL any NAF weight <= 3` FIRED | M1 1.145x, M2 1.333x | FIRED over 1.1x |
| `plant mul1` (one word `MUL = 1`) | `MUL any = 1` FIRED | M1 1.023x, M2 1.067x | flagged, under the gate: one word of 16 |
| `plant alleq` (`ROT` all 7) | `ROT all equal` FIRED | M1 0.978x (0 ops moved) | flagged; diffusion in section 6 |
| `plant rc0`, `plant rcrk0` | `RC any = 0`, `RC + rk = 0` FIRED | M1 0.978x (0 ops moved) | flagged |
## 6. Numbers
### 6.1 Classes over 2^24 days (`census-2p24.md`), with the 2^28 count (`census-2p28.md`)
Expected per day is analytic (independent draws); the NAF rows come from the exact table of `expect-231.log`.
| Class | Count 2^24 | Fraction | Expected per day | Expected count 2^24 | Count 2^28 | Worst member (day, M1 cost, M1 gain, M2 gain) |
|---|---|---|---|---|---|---|
| ROT all equal | 0 | 0 | 3.64e-11 (31^-7) | 0.001 | 0 | none |
| ROT distinct <= 3 | 534 | 3.18e-5 | 3.07e-5 | 515 | 8,229 | 2^28: day 49,986,853, 206, 1.121x, 1.000x |
| ROT distinct <= 4 | 26,010 | 1.55e-3 | 1.54e-3 | 25,783 | 412,698 | 2^28: day 208,103,482, 197, 1.173x, 1.000x |
| ROT max multiplicity >= 4 | 35,631 | 2.12e-3 | 2.35e-3 (first order) | 39,421 | 568,423 | 2^28: day 115,569,197, 200, 1.155x, 1.000x |
| ROT same-word pair sums to 32 | 2,062,481 | 0.1229 | 0.1229 | 2,062,288 | 32,997,484 | 2^28: day 97,502,921, 196, 1.179x, 1.000x |
| ROT any pair sums to 32 | 10,022,037 | 0.5974 | 0.6007 (approx., pairs not independent) | 10,078,561 | 160,353,891 | 2^28: day 27,952,752, 194, 1.191x, 1.000x |
| ROT all 8 in {1, 2, 30, 31} | 2 | 1.19e-7 | 7.68e-8 | 1.29 | 19 | day 14,330,190, 217, 1.064x, 1.000x |
| ROT >= 6 in {1, 2, 30, 31} | 1,761 | 1.05e-4 | 1.02e-4 | 1,716 | 27,651 | 2^28: day 181,528,254, 204, 1.132x, 1.000x |
| ROT >= 4 in {8, 16, 24} | 74,541 | 4.44e-3 | 4.46e-3 | 74,756 | 1,196,376 | day 5,517,722, 198, 1.167x, 1.000x |
| MUL any = 1 | 0 | 0 | 7.45e-9 | 0.125 | 4 | 2^28: day 196,441,106, 221, 1.045x, 1.067x |
| MUL any = 2^32 - 1 | 0 | 0 | 7.45e-9 | 0.125 | 1 | 2^28: day 39,988,645, 215, 1.074x, 1.067x |
| MUL any popcount <= 2 | 0 | 0 | 2.38e-7 | 4.0 | 57 | 2^28: day 218,029,468, 209, 1.105x, 1.067x |
| MUL any popcount <= 4 | 612 | 3.65e-5 | 3.72e-5 | 624 | 10,132 | day 7,275,755, 200, 1.155x, 1.000x |
| MUL any NAF weight <= 2 | 4 | 2.38e-7 | 4.62e-7 | 7.75 | 125 | 2^28: day 63,704,833, 205, 1.127x, 1.067x |
| MUL any NAF weight <= 3 | 216 | 1.29e-5 | 1.30e-5 | 218 | 3,515 | 2^28: day 247,161,685, 200, 1.155x, 1.067x |
| MUL any NAF weight <= 4 | 3,637 | 2.17e-4 | 2.20e-4 | 3,683 | 58,667 | 2^28: day 81,133,010, 198, 1.167x, 1.000x |
| MUL any < 256 | 22 | 1.31e-6 | 9.54e-7 | 16 | 262 | 2^28: day 241,187,962, 203, 1.138x, 1.067x |
| MUL two equal | 0 | 0 | 5.59e-8 | 0.94 | 18 | 2^28: day 223,900,428, 226, 1.022x, 1.000x |
| MUL M2 k >= 2 (gain >= 1.143x) | 0 | 0 | 7.9e-11 (C(16,2) x (8.12e-7)^2, approx.) | 0.0013 | 0 | none |
| RC any = 0 | 0 | 0 | 3.73e-9 | 0.062 | 1 | 2^28: day 109,542,046, 243, 0.951x, 1.000x |
| RC any popcount <= 4 or >= 28 | 5,186 | 3.09e-4 | 3.09e-4 | 5,180 | 82,804 | 2^28: day 53,303,116, 206, 1.121x, 1.000x |
| RC + rk = 0 for any of the 72 keys | 10 | 5.96e-7 | 2.68e-7 | 4.5 | 87 | day 3,194,363, 218, 1.060x, 1.000x |
| RC two equal | 1 | 5.96e-8 | 2.79e-8 | 0.47 | 8 | 2^28: day 182,857,055, 222, 1.040x, 1.000x |
Every class sits at its expectation (the largest deviation, "ROT max multiplicity >= 4", is against a first-order
bound). The worst member of every class owes its gain to its MUL draw (M1 is a MUL-only quantity); the class itself
moves nothing. No day in 2^28 has two words of NAF weight at most 3, so M2 never exceeds 1.067x.
### 6.2 The M1 tail: census against the exact expectation (`census-2p24.md`, `expect-231.log`)
| M1 cost per application | Gain vs median 231 | Days in 2^24 | Cumulative fraction, census | Cumulative fraction, exact |
|---|---|---|---|---|
| 197 (the 2^24 minimum, day 4,819,563) | 1.173x | 1 | 5.96e-8 | 8.18e-8 |
| 200 | 1.155x | 10 | 8.34e-7 | 8.62e-7 |
| 205 | 1.127x | 250 | 2.94e-5 | 2.87e-5 |
| 208 | 1.111x | 1,382 | 1.84e-4 | 1.83e-4 |
| 209 | 1.105x | 2,387 | 3.26e-4 | 3.24e-4 |
| 210 | 1.100x | 3,887 | 5.58e-4 | 5.64e-4 |
| 231 (median) | 1.000x | 1,079,174 | 0.522 | 0.522 |
Mean cost 231.113 (exact 231.111), sd 6.190 (exact 6.190). The 2^28 minimum is 194 (1.191x, day 27,952,752). A
single NAF weight has mean 11.44 and sd 1.55 over the 2^31 odd constants (`expect-231.log`).
### 6.3 ROT structure (`census-2p24.md` histograms)
| Distinct rotation amounts a chip must wire | Days in 2^24 | Fraction | Expected S(8,d) 31_d / 31^8 |
|---|---|---|---|
| 1 | 0 | 0 | 3.63e-11 |
| 2 | 4 | 2.4e-7 | 1.4e-7 |
| 3 | 530 | 3.16e-5 | 3.05e-5 |
| 4 | 25,476 | 1.52e-3 | 1.51e-3 |
| 5 | 421,405 | 0.0251 | 0.0251 |
| 6 | 2,773,843 | 0.1653 | 0.1653 |
| 7 | 7,302,781 | 0.4353 | 0.4351 |
| 8 | 6,253,177 | 0.3727 | 0.3729 |
Small amounts {1, 2, 30, 31} and byte-aligned amounts {8, 16, 24} follow Binomial(8, 4/31) and Binomial(8, 3/31) to
within 3 percent in every bin.
### 6.4 Diffusion of the worst members (`avalanche.log`: 2,048 states x 512 input bits, mean and minimum per-output-bit flip probability)
| Day | Why | ROT | After 1 application, mean / min | After 2, mean / min |
|---|---|---|---|---|
| 20,729 | genesis, same-word pair 11 + 21 = 32 | 6 25 5 25 29 11 9 21 | 0.461 / 0.383 | 0.500 / 0.498 |
| 4,819,563 | M1 worst in 2^24 | 26 18 8 30 24 24 6 9 | 0.467 / 0.426 | 0.500 / 0.499 |
| 27,952,752 | M1 worst in 2^28 | 11 26 11 7 6 20 20 3 | 0.460 / 0.392 | 0.500 / 0.498 |
| 11,482,247 | 3 distinct amounts, multiplicity 5 | 12 19 19 4 19 12 19 19 | 0.453 / 0.374 | 0.500 / 0.499 |
| 14,330,190 | all 8 amounts in {1, 2, 30, 31} | 1 1 2 31 1 31 1 31 | 0.331 / 0.196 | 0.4995 / 0.497 |
| 332,924 | NAF weight 3 word, three amounts of 1 | 1 23 1 1 12 11 16 18 | 0.458 / 0.370 | 0.500 / 0.498 |
| 196,441,106 | `MUL = 1` word (2^28) | 30 24 23 5 11 28 12 9 | 0.461 / 0.364 | 0.500 / 0.499 |
| 109,542,046 | `RC = 0` word (2^28) | 28 5 4 31 25 28 12 4 | 0.459 / 0.348 | 0.500 / 0.499 |
| planted all 1 | the worst all-equal draw | 1 x 8 | 0.345 / 0.216 | 0.500 / 0.499 |
| planted all 16 | half-word swaps | 16 x 8 | 0.387 / 0.312 | 0.500 / 0.499 |
| planted all 7 | | 7 x 8 | 0.464 / 0.400 | 0.500 / 0.498 |
The slowest draw that can exist (all rotations by 1, probability 31^-8 per day) reaches full avalanche after 2 of the
8 applications between cache reads; the worst real day in 2^28 (all amounts in {1, 2, 30, 31}) the same. No draw
gives an attacker a shorter dependency between reads than the round margin F2 measures.
### 6.5 The chain's first 100 years (`census-100y.md`: days 20,729 to 57,253 under the interim day rule)
| Item | Value |
|---|---|
| Days over 1.1x under M1 | 6 of 36,525 (1.64e-4; the 2^24 rate predicts 12) |
| First such day | 22,633 (genesis + 1,904 days, about 5.2 years in), cost 208, 1.111x |
| Worst day | 29,337 (genesis + 8,608 days, about 23.6 years in), cost 206, 1.121x |
| Days at exactly 1.100x (cost 210) | 6 more: 25,605; 28,102; 31,573; 33,710; 42,573; 54,884 |
| M2 k >= 2 | 0 |
| Genesis day 20,729 | cost 226, 0.978x; the next four devnet days (20,730 to 20,733) read 0.987x, 1.036x, 0.947x, 0.979x |
| Rotation structure | 1 day with 2 distinct amounts (57,146, genesis + 36,417, cost 225, 1.027x), 46 with 4, none with 3 or fewer otherwise; no day with a `MUL` of NAF weight under 4 |
### 6.6 Verifier time (exclusive hold, `timing.log`)
A confirmation row only: the verifier's code path is value-independent, so the exact metric is the op count above
and a wall-time difference between days can only be noise. Queued on the box at 09:47 UK (`timing.sh`, nohup, an
exclusive `flock -x -w 7200` behind the shared holds of F1, F2, F8, F9, F10 and F7 and the queued exclusive hold of
F6; the first attempt, queued 09:14 UK, was attached to a Mac ssh session and was withdrawn in favour of the nohup
job). Cores 16 and 64, nice 19, `--warps 100`, day 20,729 against day 4,819,563 (the 2^24 M1 worst), A B A B.
| Day | Cold warp 0 (ms) | Average per warp, 100 warps (ms) |
|---|---|---|
| 20,729 (genesis) | pending (`timing.log`) | pending |
| 4,819,563 (M1 worst, 1.173x) | pending (`timing.log`) | pending |
The verdict does not rest on this row.
### 6.7 The worst days in full (`worst-days.log`, `export-29337.log`)
The worst day in adders per application against the census median, in each set. M1 is the sum of the 16 NAF weights
less 16 plus 64. Every one is an ordinary draw whose 16 weights happen to sum low; none has a word under NAF weight 7.
| Set | Chain day | Years after genesis | ROT | MUL words (hex) | NAF weights | M1 cost | Gain vs median 231 | M2 |
|---|---|---|---|---|---|---|---|---|
| The public calendar, first 36,525 days (what an auditor runs) | 29,337 | 23.6 | 24 12 18 11 14 26 29 21 | 3fe4d03b 227c2043 06011627 40c10137 00234d99 063071d9 91e5abb7 035240b1 f40bfe47 809251b9 1ce999ef 940b381d da13a021 f75f8ba7 3f59bca7 01310e05 | 9 8 9 8 10 10 12 10 9 10 12 11 10 11 11 8 (sum 158) | 206 | 1.121x | 1.000x |
| 2^24 (the gate census) | 4,819,563 | 13,139 | 26 18 8 30 24 24 6 9 | a0653c83 a09de525 810085fb 6a00eba1 bf8205ff bba82079 f27da4c3 2cb80223 6001efcf 1c2814f7 ae9d09d7 ffedd7b7 943dde01 39ff47e1 0513a83f c028eef9 | 11 11 7 10 7 10 12 10 7 9 13 8 8 8 9 9 (sum 149) | 197 | 1.173x | 1.000x |
| 2^28 (extended) | 27,952,752 | 76,481 | 11 26 11 7 6 20 20 3 | f15eb273 227a08f1 20f822e1 6d477779 8d9b3aff 03040503 27fff521 bfd9ce7d 7708000d 5d60ba11 2d40005b f07e10d7 1deefdb1 4881e821 01e1fc71 3ee7c39b | 13 9 8 11 11 7 7 11 7 11 9 9 8 8 7 10 (sum 146) | 194 | 1.191x | 1.000x |
Reproduction, through the harness: `attack-f4 day --index 29337 --median 231` (and 4819563, 27952752). Through
`igneum-pow` itself, with the day bytes `"igneum-day/" || d_le64` as hex (day 29,337 = 0x7299):
`igneum-pow export --seed x --epoch-hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 --day-hex 69676e65756d2d6461792f9972000000000000 --program-class v4 --out <dir>`
writes the day's constants into the pack's `memhard.h` as `IGNEUM_MIX_ROT_INIT` and `IGNEUM_MIX_MUL_INIT`; run on the
box at 09:52 UK (`export-29337.log`, OVERALL PASS, cache FNV-1a 64 `1979492fb76b52ce`), the pack's 8 rotations and
16 multipliers equal the harness's word for word. The day-hex strings of the other two days are in section 6.2's
source list (`census-2p24.md` and `census-2p28.md`, "The 16 lowest-cost days"): `...2f6b8a490000000000` and
`...2f7086aa0100000000`.
## 7. Gate line and consequences
Gate (plan 1.4 (3)): the fraction of days with any gain over 1.1x under 2^-20.
| Model | Fraction over 1.1x | Gate | What the number means per tier |
|---|---|---|---|
| M1 (per-day LUT bitstream) | 3.26e-4 (1 day in 3,070; 2^24 and 2^28 agree; exact expectation 3.24e-4) | FAIL by 342x | An FPGA farm that re-synthesises its bitstream every day gains 10 to 19 percent on those days: 3.26e-4 x about 0.12 = 4e-5 of a year's hashes, 0.004 percent. The FPGA lane is already behind every GPU tier on reads per watt (F5: 10 to 20 M reads/s/W against the gate's 27 M), so no home miner (8, 12, 16, 24 or 32 GB), rig or pool on any vendor or OS sees a competitor appear, and no day's difficulty moves by a measurable amount |
| M2 (DSP-bound FPGA) | 0 in 2^28 | PASS | nothing moves for any tier |
| Chip (programmable constants, the chip-model-v3 recompute chip) | 0 by construction | PASS | nothing moves; a taped-out chip cannot specialise per day |
| GPU and the CPU verifier | 0 by construction | PASS | every tier pays the same ops on every day |
What the worst day buys, priced for the per-day LUT datapath (the M1 attacker) on the worst calendar day, 29,337:
| Item | Value | Source |
|---|---|---|
| Fewer adders per mixer application that day | 231 to 206, 10.8 percent fewer | section 6.7 |
| Item derivations per unit of fabric that day | 1.121x (M1 gain) | section 6.7 |
| Hash rate of a recompute FPGA (items derived per hash, the `chip-model-v3` ops-per-hash attacker) that day | up to 12.1 percent above its ordinary day, an upper bound: the 128 dependent cache reads per hash and the shadow block are untouched by the draw, so the whole-hash gain is below the mixer's | `chip-model-v3.md` section 1 (ops per hash = 128 x 72 x 130); section 3 |
| Hash rate of the stored-dataset (f = 1) FPGA or chip that day | 0 (it derives no items per hash; the mixer is paid once in the daily build) | `funding.md` B2 rank 2 |
| Days a century at or over 1.1x | 12 (6 over, 6 at exactly 1.100x) | section 6.5 |
| Share of a century's hashes the M1 attacker gains | 12 / 36,525 x about 0.11 = 3.6e-5, 0.004 percent | arithmetic on the rows above |
| What one bitstream a day costs | one place-and-route of a large part: 42 to 160 minutes on a mid-size part (PRflow, FPT 2019, cited in spec 01 section 1.13), hours on a large one; on a rented 96-thread box (Hetzner AX162 class, about USD 0.35 per hour, approximate) under USD 3 per bitstream (approximate), and it compiles any time ahead because the calendar is public | spec 01 section 1.13; price approximate |
So the bitstream is cheap and the gain is 0.004 percent of a century for the slower of the two FPGA designs: nothing
a home miner on any card, a rig or a pool on any vendor or OS can see, and nothing that moves a day's difficulty.
What is being done about the M1 line: class v4 is not changed (it is the object on the live devnet's vote). The
rejection-and-redraw rule of section 8 is proposed to main for the next class, unless main reads the census as a
fault beyond the metric (this row does not: every class sits at its expectation and the worst day is an ordinary
draw). The rule costs one redraw on 5.6e-4 of days, changes no existing vector (day 20,729 has NAF sum 178; the first
day the rule would redraw is 22,633, about 5.2 years after genesis, section 6.5), and makes the M1 gate pass by
construction. `igneum-pow` is untouched by this row.
## 8. Proposed fix for the next class: a rejection-and-redraw rule on the MUL draw (for main's decision)
Shape, like the program acceptance rule 1.4.6 and `DeriveProgram::check`: draw the 16 `MUL`, test, and on rejection
continue the same stream with 16 fresh draws (so every later draw keeps its position only within an accepted block;
`RC` is drawn after the accepted `MUL` block). Tests, in order:
| Rule | Threshold | Rejection probability per candidate | What it closes |
|---|---|---|---|
| Sum of NAF weights of the 16 `MUL` at least 163 (M1 cost at least 211, gain at most 1.095x against the median 231) | `sum_i NAF(MUL_i) >= 163` | 5.64e-4 (`expect-231.log` cumulative at cost 210) | the M1 tail: no day over 1.1x by construction |
| Every `MUL` of NAF weight at least 4 | `NAF(MUL_i) >= 4` | 1.30e-5 per day | `MUL = 1`, `2^32 - 1`, `2^a +- 1`, `2^a +- 2^b +- 1`: the M2 words (hygiene; M2 already passes) |
| `ROT`: at least 4 distinct amounts (the `DISTINCT_ROTS_FLOOR` idea of `derive.rs`) | `distinct >= 4` | 3.07e-5 per day | the degenerate rotation draws (hygiene; 0 ops moved, diffusion fine at 2 applications) |
Total rejection about 6.1e-4 per day: one redraw every 4.5 years of chain time; `MAX_ATTEMPTS`-style exhaustion is
impossible in practice (64 rejections in a row at 6e-4 each). Class check to land with it: a unit test in `memhard.rs`
that plants a low-sum draw (stream seed chosen so the first MUL block fails) and asserts the redraw, plus this
harness re-run over 2^24 showing 0 days over 1.1x under M1 after the rule. The reproduction line for the finding
without the rule: `attack-f4 day --index 4819563 --median 231` (cost 197, 1.173x) and
`attack-f4 day --index 27952752 --median 231` (cost 194, 1.191x).
Consensus consequence: the rule changes the day-key-to-constants map on rejected days only, so it must land before the
freeze tag. In the chain's first 100 years the sum rule redraws 12 days (6 under 1.1x and 6 at exactly 1.100x,
section 6.5), the first of them 22,633, about 5.2 years after genesis; no pack cut for the devnet, the testnet or the
first five years of mainnet changes. The per-word rule redraws no day in the first 100 years (no word of NAF weight
under 4 in `census-100y.md`); the `ROT` rule redraws one, day 57,146 (2 distinct amounts, 99.7 years in).
## 9. What this row did not do
* It did not time the verifier per day beyond the confirmation row of 6.6: the verifier's code path is
value-independent (no branch on a drawn value), so op counts are the exact metric.
* It did not search optimal single-constant multiplication costs (not computable at 2^28 scale); NAF is the standard
canonical bound and the ratio between days is what the gate asks.
* It did not census the era draw (F7) or the spec's intent for the 64-bit seeding (F7); the fact is stated in section 1.
## 9. Reconciliation with adv-mixer-2's independent census (7 October 2026, 21:5x UK; main's order through the
crypto-engage coordinator: the two tables side by side, the median each used, one agreed figure)
Both censuses read the same class, the FPGA LUT multiplier-adder area (no chip or GPU gain), with the same cost form,
64 + the sum over the sixteen MUL of (signed-digit weight - 1). They differ in one convention. This harness's
`naf_weight` (src/main.rs line 72) computes the canonical NAF of the constant as a `u64` and counts every digit,
including the carry digit at position 32 that a 32-bit odd constant carries with probability 1/3; adv-mixer-2's
`w32` counts the digits at positions 0 to 31 only, because a multiplier modulo 2^32 is never built with a digit at
position 32. Checked on 2^18 random odd constants: mean weight 11.442 with the carry digit against 11.109 without
(P(digit at 32) = 0.3334), cost 231.1 against 225.7. So this record charged every day about 5 adders a 32-bit
multiplier never pays, unevenly per day, which is the whole of 231 against 226 and of 12 against 15 days a century.
| Quantity | This record (M1, carry digit counted) | adv-mixer-2 (model A, w32 modulo 2^32) | Agreed |
|---|---|---|---|
| Cost form | 64 + sum(NAF(MUL_i) - 1), NAF over u64 | 64 + sum(w32(MUL_i) - 1), digits 0 to 31 | adv-mixer-2's: the digit at position 32 costs nothing in a 32-bit multiply |
| Median per application | 231 (census over 2^24 chain days; mean 231.113) | 226 (exact, 16-fold convolution over all 2^31 odd constants; the log's provisional 231 superseded) | 226 |
| Threshold for a 1.1x gain | cost under 210 (NAF sum under 163) | cost A at most 205 | cost A at most 205 |
| Fraction of days over 1.1x | 3.264e-4 (5,476 of 2^24); exact expectation 3.243e-4 | 5.677e-4 (2^24 census); exact 5.694e-4 | 5.69e-4, about 2^-10.8, against the 2^-20 gate |
| Days over 1.1x per century (36,525 public days) | 12 | 15, listed in its Q4 | 15 |
| Worst public day | 29,337 (2050-04-28), cost 206, 1.121x | 29,337 (2050-04-28), cost A 203, 1.113x | 29,337, 1.113x |
| Genesis day 20,729 | cost 226, 0.978x | cost A 219, 1.032x | 219 |
| DSP-bound metric | M2: 0 days with k >= 2 in 2^28; k >= 1 at 1.067x | model C: P(k >= 1) 3.12e-5 (1.067x), P(k >= 2) 4.57e-10 (1.143x) | agree: under the gate at every k |
| Redraw rule for the next class | reject NAF sum under 163 (cost under 211), redraw from the next stream values | cost A at most 205, or k >= 1, or the eight ROT equal: continue the same SplitMix64 stream and draw the forty again | adv-mixer-2's form and numbers |
Verdict unchanged: PASS against class v4 on the DSP-bound reading (M2 and model C agree at 0), the LUT tail bounded
and measured; the agreed numbers above replace this record's M1 figures wherever quoted, and AP-F4-1's rule on the
v5 list takes adv-mixer-2's form. Harness note: `attack-f4 plant` and `day` default `--median` to 221 when the flag is
absent (the plant lines of the class v5 run this evening read "median 221" for that reason); every census line in
this record was run with `--median 231` and is restated against 226 above.

View file

@ -0,0 +1,229 @@
# F6: the verifier's worst case over 10^5 class v4 programs
Attack-pass row F6 (`docs/plans/cryptanalysis.md` 4.2; the record `docs/analysis/attack-pass-2026-10.md`), the box search.
The O-1.14 laptop relay run is not in this record (main runs it separately). Written 7 October 2026.
## Target
| Item | Value |
|---|---|
| Commit | `924288d1` on branch `attack-pass` (`igneum-pow` is byte-identical at the worktree HEAD `8e36faf6`: `git diff --stat 924288d1..HEAD -- igneum-pow` is empty) |
| Class | `--program-class v4`: generator 4 on `V4_CLASS` = `mx8+sh256x27` (`LoadClass::MX8` plus `ShadowClass { instrs: 256, reps: 27 }`), no era bytes (the same draw `igneum-pow bench --program-class v4 --seed S` makes) |
| Dataset | day `2026-10-03`, `Shape::for_class_day(V4_CLASS, 0)`: cache 2^26 words (256 MiB), mixer x8, dataset 2^28 words, memory-hard |
| Work per hash | 64 base instructions x 8 iterations (16 loads) plus 256 shadow instructions x 27 passes x 8 iterations = 55,296 shadow instructions, 101,192 counted ops at the 1.83 convention |
| Gate | 10 ms per 32-lane warp, cold, on the half-core proxy (plan 1.4 item 6; spec 01 section 1.9 and 1.11; `algorithm.md` 3.3 and 5.5) |
| Programs | 10^5 deterministic string seeds `attack-f6/0` to `attack-f6/99999` through the class v4 chain draw with its acceptance rule (5.22 percent needed a second or third attempt, max attempt 3) |
The era draw is not in the search: `generator.rs` draws the shadow block from `NONLOAD_WEIGHTS` with no era perturbation
(no `perturb` path exists in the code at this commit), and the era parameters change only the load addressing, not the op
counts. Every drawn program has exactly 48 non-load base instructions and 256 shadow instructions, so the verifier's cost
differs between programs only through the family mix (the per-family cost on the CPU) and the data.
## Known-failed shape
A drawn program whose verifier warp exceeds 10 ms cold on the half-core proxy. The acceptance rule (spec 1.4.6) bounds the
miner's side (distinctness, bias, saturation); nothing bounds the verifier's cost per program, and the half-core headroom
of the average program is 1.8 ms (`algorithm.md` 5.5), so a family mix that costs the CPU interpreter more than the average
could cross the gate.
## Harness
`tools/attack/f6-verifier/` (its own cargo crate, `igneum-pow` as a path dependency, the same release profile as the CLI:
opt-level 3, LTO, one codegen unit). It builds the day's dataset ONCE (`DatasetSource::new_shape`, 0.58 s on the box) and
swaps programs under it: `Epoch { program, dataset }` is only the pair, and `verify::hash_warp(&program, base, &dataset)`
takes both, so one dataset serves every program. The naive path (`igneum-pow bench` per seed) refills the cache every
time (370 ms) and would take 10 hours per core.
| Command | What it does |
|---|---|
| `attack-f6 scan --program-class v4 --count N --start S --threads T --cold-reps R --flush swap --out F` | program i = seed `attack-f6/<S+i>`; one CSV line per program: generation time, R timed warps (each after a flush), their min, the op counts by family for the base and the shadow block |
| `attack-f6 time (--program-class v4 \| --class dr736) --seeds-file F --cold-reps R --steady W --flush sweep` | the deep re-time: R cold warps (each after a 256 MiB write sweep, what the cache fill does before `bench`'s "single cold run"), max, median, min, a steady average of W warps, and `GATE 10 ms PASS/FAIL` on the max |
| `attack-f6 micro --program-class v4` | the genesis program with its shadow block rewritten to one family at a time against the same program with no shadow: the per-family cost of a shadow instruction (ranking weights only) |
| `attack-f6 load --program-class v4 --seconds 0` | hashes class v4 warps on the calling core until killed: the SMT sibling's load for the half-core proxy, the same class the 3.3 proxy ran on both siblings |
| `rank.py --scan ... --weights ... --column ... --top 50 --out-prefix P` | the proxy ranking (sum over families of weight x (8 x base count + 216 x shadow count)), the distribution (min, median, p99, p99.9, max with the seed), the worst-N lists, a no-intercept regression of time on the family counts |
Build line (from the crate directory):
`IGNEUM_AGENT=attack-f6 bash /Users/joshm/Projects/igneum/tools/build-remote.sh --artefacts "target/release/attack-f6" --out <scratch> -- build --release`
(build 08:04 to 08:06 UTC, rustc 1.99.0, binary sha256 `89c35674...17f3f9`, on the box at
`/srv/builds/igneum-wt-attack/tools/attack/f6-verifier/target/release/attack-f6`).
Box scratch: `/srv/builds/igneum-wt-attack/target-attack-f6/` (not the `attack-f6/` the brief named: that path is an
untracked directory in the worktree mirror and `build-remote.sh`'s checkout step runs `git clean -fd` on the mirror
before every build of this worktree by any agent, which removed it once at 08:04 UTC; `target-*` is on the clean's keep
list, so this name survives). Logs there: `phase1.log`, `scan-0.csv`, `scan-50000.csv` (phase 1), `phase2a.log`,
`clock-2a.log`, `full-0.csv`, `phase2b.log`, `clock-2b.log`, `full-50000.csv`, `phase2c.log`, `clock-2c.log` (phase 2),
`smoke.log` (the functional check). Copies on the Mac under the session scratchpad `attack-f6/`.
Run lines:
| Phase | Hold | Cores | Line |
|---|---|---|---|
| 1, pre-screen (counts and a coarse time, no timing claim) | `flock -s` per 50,000-program chunk (2.6 min each) | `nice -n 10 taskset -c 42-47,90-95`, 12 threads | `attack-f6 scan --program-class v4 --count 50000 --start {0,50000} --threads 12 --cold-reps 2 --flush swap` |
| 2A, firings, micro, full pass first half | `flock -x` | `nice -n 19 taskset -c 40`; half-core: core 88 running `attack-f6 load` | `phase2a.sh` |
| 2B, full pass second half, worst 50 by proxy and worst 50 by coarse time re-timed on both proxies | `flock -x` | same | `phase2b.sh` |
| 2C, worst 1,000 by the full one-core pass on the half-core; worst 10 deep re-timed on both proxies | `flock -x` | same | `phase2c.sh` |
The clock of cores 40 and 88 (`scaling_cur_freq`) and the load average were read every 5 s during every exclusive hold
(`clock-2*.log`).
## The two firings (batch A, exclusive hold taken 08:15:20 UTC, core 40 at 3,799.9 MHz throughout, `clock-2a.log`)
`attack-f6 time ... --cold-reps 5 --steady 20 --flush sweep`, the gate applied to the max of the 5 cold warps
(`phase2a.log`). Lane-0 vectors equal the known ones (dr736 `e23d389f3eea0c83`, the Mac's).
| Case | Proxy | Cold max / median / min (ms) | Steady, avg of 20 | Known value (`algorithm.md` 3.3) | Verdict |
|---|---|---|---|---|---|
| known-fail, `--class dr736`, genesis seed | one core | 10.284 / 9.982 / 9.952 | 9.562 | 10.51 cold, 9.76 steady | FAIL (fired) |
| known-fail, `--class dr736`, genesis seed | half-core | 14.135 / 13.795 / 13.400 | 13.225 | 15.49 | FAIL (fired) |
| known-pass, `--program-class v4`, genesis seed | one core | 5.156 / 5.140 / 5.135 | 4.909 | 5.06 cold, 4.90 steady | PASS (fired) |
| known-pass, `--program-class v4`, genesis seed | half-core | 8.624 / 8.510 / 8.389 | 8.268 | 8.23 | PASS (fired) |
The harness reads the known-fail class over the gate and the known-pass class under it on both proxies, within 2 percent
of the 3.3 one-core numbers and within 9 percent on the half-core (the earlier half-core run loaded the sibling with
`igneum-pow bench` of the same class; this one hashes class v4 warps on it continuously).
## What one program can move (batch A, `micro`, core 40 solo)
The genesis class v4 program with no shadow block: 4.511 ms steady; with its drawn shadow: 4.920 ms. So the whole 55,296-
instruction shadow block costs 0.41 ms per warp on this core (7.4 us per 1,000 shadow instructions; `model.py` carries
7.0) and the base program with its 128 loads and 4,096 item derivations costs the other 4.5 ms. The verifier's cost is
92 percent dataset derivation (the x8 mixer, 8 dependent cache reads per item), which no drawn program changes: every
class v4 program has 16 loads and the acceptance rule's distinctness test keeps the items per warp near 4,096. The family
mix of the shadow can move at most a fraction of 0.41 ms. The per-family rewrite (the 256 shadow instructions all one
family) reads add 4.798, sub 4.703, xor 4.683, rotl 4.818, mad 4.805, shfl 5.042, rotr 5.160 ms per warp; the mul,
mulhi and or rows (0.97, 1.55, 1.13 ms) are degenerate (the registers collapse to 0 or all-ones, every lane then loads
the same item and the memory side vanishes) and are not ALU costs. Batch A's micro line printed its per-1,000 column
1,000x too small (ns per instruction); fixed in the source, the numbers above are the ms-per-warp column, which is right.
## Full one-core pass A (batch A, 50,000 programs, core 40 solo, one cold warp each after a program swap, `full-0.csv`)
617 s for 50,000 programs (12.3 ms each, 2.5 ms of it generation and acceptance). The per-family regression on the
50,000 timings (no intercept) gives 86 to 98 us per 1,000 executed instructions by family, R^2 0.001: the family mix
explains none of the program-to-program variation. Those regression weights (add 89.1, sub 87.7, mul 90.6, mulhi 98.4,
xor 89.6, or 86.0, rotl 89.4, rotr 86.3, mad 91.2, shfl 95.6) are the proxy weights used for the worst-50-by-proxy list.
| Pass A | n | min | median | p99 | p99.9 | max |
|---|---|---|---|---|---|---|
| all rows | 50,000 | 4.610 (`attack-f6/1278`) | 4.950 | 6.753 | 7.834 | 10.710 (`attack-f6/26705`) |
| rows outside the three disturbed blocks | 44,000 | 4.610 | 4.948 | 5.606 | 5.662 | 6.194 (`attack-f6/48484`) |
The rows over 6 ms sit in three 2,000-program blocks (26,000 to 27,999: 337 rows; 36,000 to 37,999: 300; 38,000 to
39,999: 294) and nowhere else (0 in each of the other 22 blocks); the neighbours of the 10.71 ms seed, unrelated programs,
all read 7.6 to 8.5 ms. `clock-2a.log` shows core 88 (the idle sibling of a solo run) at 3.8 GHz for stretches in those
minutes (08:21:25, 08:21:45, 08:22:05 to 08:22:25 UTC) and core 40 dipping to 3.68 GHz at 08:21:00: a foreign process
(the box's hands run outside the measure lock) sat on the sibling, which is the half-core condition, and those rows read
half-core numbers. They are not program properties and not numbers; the three blocks are re-scanned in batch B, and from
batch B on a per-second sampler logs every process whose last CPU was 40 or 88 (`clock-2b.log`, `clock-2c.log`) so a
disturbed row can be named.
## Batches B and C: queued, starved of the exclusive hold (state at 09:46 UTC)
Batch B (the second 50,000 of the full one-core pass, the re-scan of the three disturbed blocks, the worst 50 by proxy
and the worst 50 by phase-1 coarse time re-timed 10 cold reps each on both proxies) was queued with `flock -x -w 7200`
at 08:33:56 UTC and had not taken the file by 09:46 UTC. Linux `flock` gives a pending exclusive waiter no priority over
new shared takers; eight lanes re-take the file in chunks (17 to 38 shared holders at every reading, F2 spawning many
short ones, one F9 hold 33 minutes old at 09:06, over the 30-minute cap), so the file is never free. Left on the box:
`run2b-retry.sh` re-queues batch B up to four more times (2 h each); `run2c-auto.sh` waits for `BATCH B DONE`, builds the
batch C lists on the box (`mklists.py`: the worst 1,000 and worst 10 by the full one-core pass, pass A's three disturbed
blocks replaced by their re-scan) and queues batch C (the worst 1,000 on the half-core at 2 cold reps each, the worst 10
at 20 cold reps on both proxies). A marker sits beside the lock (`/srv/builds/_locks/measure.wanted-by-attack-f6`). When
they land, `timeparse.py --log phase2b.log --clock clock-2b.log` (and `2c`) prints the per-seed tables with the sampler's
foreign-process column, and this record is completed.
## Numbers so far, the gate, the verdict
| Quantity | Value | Where |
|---|---|---|
| Programs drawn and counted (phase 1) | 100,000 | `scan-0.csv`, `scan-50000.csv` |
| Programs timed cold on core 40 alone (one-core proxy) | 50,000 (44,000 clean, 6,000 in disturbed blocks awaiting the re-scan) | `full-0.csv` |
| One-core cold, clean rows: min / median / p99 / p99.9 / max | 4.610 / 4.948 / 5.606 / 5.662 / 6.194 ms (`attack-f6/48484`), single warps, not yet re-timed | `full-0.csv` |
| Genesis class v4, half-core, max of 5 cold | 8.624 ms | `phase2a.log` |
| Half-core over one-core, genesis class v4 | 1.67x (8.624 / 5.156) | `phase2a.log` |
| Programs timed on the half-core proxy | 1 (the genesis seed) | `phase2a.log` |
| Core 40 clock during every exclusive timing | 3,799.9 MHz (dips to 3,680 MHz only in the disturbed minutes) | `clock-2a.log` |
Gate line: the worst program under 10 ms cold on the half-core proxy. Not yet measured: no drawn program other than the
genesis seed has a half-core number, and the one-core worst (6.194 ms, a single warp with no sampler running) has not been
re-timed. Carried to the half-core at the genesis ratio it would read 6.19 x 1.67 = 10.3 ms, over the gate; carried at the
additive half-core cost of the genesis program (8.624 - 5.156 = 3.47 ms) it would read 9.66 ms, under it by 0.34 ms. The
2.5x bracket of `algorithm.md` 5.5 is between. Whether `attack-f6/48484` (and the other clean rows over 5.6 ms: 1 in 100
of the pass) is a program property or a short disturbance is what batch C's 20-rep re-time with the sampler decides;
the record says FINDING if its half-core cold max reads 10 ms or more.
Verdict: INCOMPLETE. 100,000 programs drawn and ranked, 50,000 timed on the one-core proxy, 0 of the worst re-timed on
the half-core proxy; the two firings fired; the box search's second half and the half-core re-times are queued and
starved of the exclusive hold.
## The ladder ceiling implied so far
`algorithm.md` 5.5 and `model.py --section ladder` set the ceiling from the half-core headroom at 12.1 us per 1,000 shadow
instructions, N = 101,192 + instructions x 1.83.
| Worst program on the half-core | Headroom to 10 ms | Shadow instructions it buys | Ceiling N (counted ops) |
|---|---|---|---|
| 8.23 ms (3.3's average, the published figure) | 1.77 ms | 146,000 | about 370,000 |
| 8.62 ms (genesis seed, this run's max of 5) | 1.38 ms | 114,000 | about 310,000 |
| 9.66 ms (48484 if the additive carry holds) | 0.34 ms | 28,000 | about 152,000 |
| 10.3 ms (48484 if the 1.67x carry holds) | none | 0 | below today's 101,192: the floor rung 100,000 is the ceiling |
The proposed genesis ladder {100,000; 130,000; 200,000; 330,000; 650,000; 1,000,000} already exceeds the 310,000 ceiling
at its fourth rung on the genesis program alone; on the worst program the ceiling could be the floor. This is the
ladder's own open question (plan 1.1, "ceiling set by the verifier"), and the number that sets it is the half-core
worst case still queued.
## Consequences per user tier (at the numbers measured so far; the model's table of 5.5 at 10 ms beside them)
| | At 8.62 ms (genesis, half-core max) | At 9.66 ms (worst, additive carry, unverified) | At 10.3 ms (worst, 1.67x carry, unverified) | Model at 10 ms |
|---|---|---|---|---|
| A node on a 2019-class laptop core at 1 bps | 0.9 percent of one core | 1.0 | 1.0 | 1 |
| At 10 bps (the Devnet 2 experiment) | 8.6 percent of one core | 9.7 | 10.3 | 10 |
| IBD over the 108,000-header pruning window, one core | 15.5 min | 17.4 | 18.5 | 18 |
| Header flood: invalid headers per second that saturate one core | 116 | 104 | 97 | 100 |
| A pool core verifying shares, shares per second per core | 116 | 104 | 97 | 100 |
What each tier does with it: a home miner (8, 12, 16, 24 or 32 GB card, any vendor, any OS) runs a node that spends about
1 percent of one CPU core on the hash at 1 bps whatever the drawn program, and 9 to 10 percent at 10 bps; the card is not
involved. A rig is the same per node. A pool verifying shares at 100 per second per core needs one core per 100 shares
per second at the worst program, 116 at the average: a pool that sized its share verification at the average loses 14
percent of its per-core headroom on the worst program, so pools size at 97 shares per second per core (the 10 ms figure)
and never at the average. A node under header flood holds at about 100 invalid headers per second per core on any
program, the M15 figure. The 2019-class core itself is still the half-core proxy until the O-1.14 laptop run lands
(main's lane).
What this lane does about it: completes batches B and C when the hold comes (automatic, on the box); if the worst
program reads 10 ms or more on the half-core proxy, the finding goes to main with the seed, the reproduction line
`igneum-pow bench --program-class v4 --seed <seed> --day 2026-10-03 --warps 50` on core 40 and the half-core, and the
proposed fix: an acceptance-rule bound on verifier cost (a per-program cost model over the family counts checked at
draw time, a redraw when it exceeds the bound, exactly as rule (c) redraws on bias) and the ladder's ceiling set from
the measured worst, not the average; `igneum-pow` is not edited by this lane.
## Batches B and C landed (7 October 2026, 13:5x UTC, cores 40 and 88 under the per-core lease)
The measure file was retired at 13:3x UTC and replaced by per-core leases; cores 40 and 88 were leased to this row
(`/srv/builds/_bin/lease cores 40,88 --owner attack-pass`), so the batches ran with nothing else on those cores while
builds continued on the rest of the box. Core 40's clock (`clock-2b.log`, `clock-2c.log`): median 3,799.9 MHz in both
batches, 954 of 958 and 57 of 63 samples at 3.7 GHz or more, the 1.5 GHz readings between runs.
| Batch | What | Programs | Reps | Worst program | Half-core cold max | Log |
|---|---|---|---|---|---|---|
| B | the second 50,000 of the one-core pass, then the worst 50 by proxy and the worst 50 by coarse time re-timed cold on both proxies | 200 re-timed | 10 | `attack-f6/87142` | 8.708 ms | `phase2b.log`, done 13:52:11Z |
| C | the worst 1,000 by the full one-core pass on the half-core, then the worst 10 at 20 cold reps on both proxies | 1,000 + 10 | 2, then 20 | `attack-f6/88521` (8.629), `attack-f6/15781` (8.414) | 8.629 ms | `phase2c.log`, done 13:57:45Z |
The one-core worst of the full pass (`attack-f6/48484`, 6.194 ms single warp) does not reach the half-core top ten:
its one-core reading was a short disturbance, as the batch C re-scan shows. Every program timed on the half-core
proxy reads under 9 ms.
## Gate line and verdict
Gate: the worst program under 10 ms cold on the half-core proxy (and on a 2019-class core: O-1.14, the i7-9700K row,
class v4 6.334 ms cold max). Result: the worst of 100,000 class v4 programs on the half-core proxy is 8.708 ms,
1.29 ms under the gate; the genesis program reads 8.624 on the same proxy, so the worst drawn program costs 1 percent
more than the genesis one and the distribution is tight (one-core clean rows 4.610 to 6.194 ms, p99 5.606).
Verdict: PASS. The two firings fired (dr736 FAIL at 15.49 ms half-core; class v4 genesis PASS at 8.62). Consequences
per tier: a 2019-class node verifying the worst class v4 program spends 0.9 percent of one core at 1 bps and 9 percent
at 10 bps on the pessimistic proxy; a header flood needs about 115 invalid headers a second to saturate one such
core; a pool verifies about 115 shares a second per core; IBD of 108,000 headers is about 16 minutes of one core.
Implied ladder ceiling on the half-core proxy from the worst program: 10 - 8.708 = 1.29 ms of headroom buys about
106,000 shadow instructions, N about 300,000 counted ops at the 1.83 convention (approximate), against 370,000
from the genesis program's headroom; the ladder's ceiling should be read from the worst program, not the genesis
one, so rung 2 (199,600) stays admissible and rung 3 (330,700) does not on this proxy.

View file

@ -0,0 +1,165 @@
# F7: the era-draw bias harness and the era-seed census
Attack-pass row F7 (`docs/plans/cryptanalysis.md` 4.2; the record `docs/analysis/attack-pass-2026-10.md`), both halves:
the fast-time re-roll harness (node lane) and the 2^20 era-seed census plus the 64-bit day-key check (hash lane).
Written 7 October 2026. Every number cites its log path on igneum-build-1.
## Target
| Item | Value |
|---|---|
| Commit | `924288d1` on branch `attack-pass` (`igneum-pow` is byte-identical at worktree HEAD `11b375a0`: `git diff --stat 924288d1..HEAD -- igneum-pow docs/spec infra/fast-time` is empty) |
| Spec | `docs/spec/01-lottery-hash.md` 1.13.1 (era seed and draw), 1.8.4 (the day-key mixer stream); `docs/spec/04-seeds-and-vdf.md` 4.4 (era seed pipeline) and 4.6 (T from a reference core); `docs/plans/era-layout.md` sections 1 and 8 (branch `ca2-era`); `docs/analysis/horizon/algorithm.md` 5.4 |
| Draw code | `igneum_pow::generator::era_draw` over `V3_ALLOWED = [1]` (the chain's path): the width draw (consumed, pinned at 4 bytes), the odd stride multiplier `M`, the rotation `R` in 1..31, the four interleave positions `pos` by partial Fisher-Yates |
| Day-key code | `igneum_pow::memhard::MixParams::with_shape`: `SplitMix64::new(K[0] \| (K[1] << 32))` draws ROT[0..7], MUL[0..15], RC[0..15]; `K = seed_words_from_bytes("igneum-day/" \|\| day_le64)` (node fork `consensus/pow/src/igneum.rs`, `bind::day_bytes`) |
| Node draw input (today) | `consensus/src/consensus/mod.rs` `seed_below`: `E_n` is the hash of the last selected-chain block below `15,552,000 n - 7,200` (era 0: genesis). The 1-hour VDF of spec 4.4 and the certified checkpoint it reads do NOT exist in the node (era-layout.md section 8, `proto-vdf` is a prototype) |
## Sub-row verdicts
| Sub-row | Verdict | Gate (plan 4.2 F7) |
|---|---|---|
| (a) re-roll harness | INCOMPLETE, with the written argument | no re-roll inside the publish window |
| (b) 2^20 era-seed census | PASS | no era class with gain over 1.1x at a fraction over 2^-20 |
| (c) 64-bit day-key seeding | PASS, within spec intent (one observation recorded) | the draw's input set as the spec states it |
## (a) The re-roll harness (node lane)
`tools/attack/f7-era/reroll.mjs`: a 3-node fast-time network (`infra/fast-time/override-60x.json` with
`skip_proof_of_work`, the `class-v4-signal.mjs` shape), own ports 29800 and up, own devnet suffix 980, own data dir
`/tmp/igneum-fast-time-attack-f7`. The node binary is the ladder fork `vendor/igneum-node-ladder` at `1591ee1d`
(`igneumd 2.1.0`, already built on the box; read-only). Two honest virtual miners share 1 block/s on nodes 0 and 1; the
adversary on node 2 holds a block `A` built on the tip at DAA score `S - 1` (the seed block sits there), optionally waits
a stub VDF of `--vdf-ms`, then publishes `A` to try to make its own block the epoch's seed block (the last selected-chain
block below the cut `S`). A re-roll succeeds when the epoch's reported seed becomes `hash(A)`.
The era cut `15,552,000 n - 7,200` is 180 days of DAA score away on every profile (`POW_ERA_BLOCKS` is a chain constant,
not an override field), so the harness attacks the EPOCH cut (`60 e - 10` at 60x), which runs the identical `seed_below`
derivation at a reachable score, one cut per minute. The harness's own era draw (JS) is checked byte-for-byte against the
Rust census at start: seed `b62532bc...` draws `M 558c0543 R 4 pos [0,1,2,3]` on both (log line "draw self-check ... OK").
Firings (both runs 6 cuts, box cores 36-37,84-85 under the shared measure lock):
| Run | `--vdf-ms` | Re-rolls to A | Gate | Harness | Log |
|---|---|---|---|---|---|
| known-pass | 0 (no delay, the stand-in) | 1 of 6 (epoch 11, seed = A) | FAIL | SOUND (fires) | `/srv/builds/igneum-wt-attack/attack-f7/reroll-knownpass.log` |
| known-fail | 5,000 (a delay past one block interval) | 0 of 6 | PASS | SOUND (silent) | `/srv/builds/igneum-wt-attack/attack-f7/reroll-knownfail.log` |
Both runs: 6 of 6 adversary blocks accepted, all three sinks agree, no reorg of the honest chain. The harness fires on the
known-pass and is silent on the known-fail, so it is trusted.
Written argument (the plan allows one for the VDF's assumptions; the VDF's own delay soundness belongs to the finality
review row of `funding.md`). The re-roll is possible ONLY when the adversary can evaluate the draw of a candidate input
inside the block publish window. Today the node has no VDF: `E_n` is a plain block hash, so the input of any candidate
block is known the instant the block is built, and the harness shows the last-block-before-the-cut is grindable with one
block of hash (1 of 6 cuts steered in fast time, `--vdf-ms 0`). With any delay past one honest block interval the
re-roll is gone (`--vdf-ms 5000`: 0 of 6). The design closes this with the 1-hour class-group VDF of spec 4.4: re-rolling
by withholding needs the 3,600 s VDF evaluated inside the 2 s window, a 1,800x evaluator, and spec 4.6's margin table
gives 300x as the horizon (`algorithm.md` 5.4; `sim/horizon/algorithm/model.py --section era`). The forge route needs
2/3 of the 30-day weight, 20 days of 100 percent hash (CLAUDE.md headline). The sub-row is INCOMPLETE because the harness
cannot demonstrate the real gate: the VDF and the certified checkpoint it reads are not in the node yet (era-layout.md
section 8 states this). What the harness DOES establish: the C_era cut rule with no delay is grindable, so the era draw's
soundness rests entirely on the VDF landing before the draw procedure is frozen, and the delay-soundness measurement is
owed to the finality lane.
## (b) The 2^20 era-seed census (hash lane)
`tools/attack/f7-era/` (a cargo crate with `igneum-pow` as a path dependency and an empty `[workspace]`; ELF built on the
box, sha256 `a87818d8...`). `attack-f7 census` runs `era_draw` over `V3_ALLOWED` on `2^n` seeds and classifies each draw;
`attack-f7 all` runs the plant known-fail case, the census, the spec-stream op-weight census and the day-key check.
Known-fail / known-pass of the classifier (planted parameters through a test hook in this crate; log
`/srv/builds/igneum-wt-attack/attack-f7/census-2p20.log`): every planted weak draw fires its flag (M = 1, M = 2^32-1,
M = 2^16+1, a naf-2 multiplier, an even M, R = 0, R = 32, pos linear, pos contiguous, pos not ascending) and a sound draw
(igneum-era-test/0) raises nothing. "Plant verdict: every planted case fired and the sound draw did not."
Census results (2^24 = 16,777,216 draws, the stronger run; `census-2p24.log`; the 2^20 run agrees, `census-2p20.log`):
| Class | Count (2^24) | Fraction | Expected (uniform) | Chip gain |
|---|---|---|---|---|
| M even (bijection failure) | 0 | 0 | 0 | finding if present: none |
| R out of 1..31 | 0 | 0 | 0 | finding if present: none |
| pos invalid (not 4 ascending) | 0 | 0 | 0 | finding if present: none |
| M = 1 (identity stride) | 0 | 0 | 4.66e-10 | 1.0034x |
| M = 2^32 - 1 | 0 | 0 | 4.66e-10 | 1.0030x |
| popcount(M) <= 2 | 1 | 5.96e-8 (2^-24) | 1.49e-8 | 1.0030x |
| popcount(M) <= 4 | 43 | 2.56e-6 (2^-18.6) | 2.33e-6 | 1.0022x |
| popcount(M) <= 6 | 1,626 | 9.69e-5 | 9.61e-5 | 1.0014x |
| popcount(M) <= 8 | 27,749 | 1.65e-3 | 1.66e-3 | 1.0007x |
| naf(M) <= 2 | 1 | 5.96e-8 | - | 1.0030x |
| naf(M) <= 3 | 18 | 1.07e-6 | - | 1.0026x |
| M = 2^k + 1 | 1 | 5.96e-8 | 1.44e-8 | 1.0030x |
| pos linear [0,1,2,3] | 9,257 | 5.52e-4 | 5.50e-4 | 1.0000x |
| pos contiguous | 120,054 | 7.16e-3 | 7.14e-3 | 1.0000x |
| pos in the low byte | 645,856 | 3.85e-2 | 3.85e-2 | 1.0000x |
The gain metric is the datapath energy a chip saves per hash against the base weights, over the hash's datapath energy
(19.5 nJ at 100,000 ops x 0.195 pJ, the N5 floor of `algorithm.md` 5.4 / `model.py --section era`). The stride multiply is
one of three address operations, run 128 times per hash (16 loads x 8 iterations); a low-weight `M` replaces the multiplier
with a few shift-adds, worth at most 128 x 0.52 pJ = 67 pJ, so M = 1 is the richest corner at 1.0034x. The rotation is a
wire mux and the interleave an address-line permute, 0 pJ on the modelled chip. No drawn parameter touches the memory
bound, the item derivation, the load count or N.
Gate: no class with gain over 1.1x at a fraction over 2^-20. The richest gain in the whole classifier is 1.0034x (M = 1),
and M = 1 did not occur in 2^24 draws (expected 4.66e-10). Every class at a fraction over 2^-20 has gain 1.0000x to
1.0007x. PASS on both counts.
Uniformity of the draw (2^24): stride rotation R over 1..31 chi-square 38.5 on 30 dof (max bucket deviation 2.07 sigma,
R = 0 or 32 seen 0 times); interleave pos 1,820 of 1,820 four-subsets seen, chi-square 1,775.7 on 1,819 dof (max deviation
3.63 sigma, 0 draws with a non-4-subset); M bit 0 always set (odd by construction), bits 1..31 each set in 0.500 of draws
(worst bit 1.81 sigma); the stride bijection never failed (0 even M). The era stream's own 64-bit seed (words 0 and 1) was
distinct on all 2^24 draws.
Op-weight corners (spec 1.13.1 first stream, implemented in `attack-f7 spec` from the spec text because `igneum-pow` does
not draw the op-weight perturbation at this commit; 2^20 draws, `census-2p20.log`): the ten non-load weights each
perturbed by -2..+2 and renormalised to 75 move the multiply share (mul+mad+mulhi, base 22 of 75) between 15 and 31. The
richest corner for a chip is 15/75 (0.152 pJ per op, -22 percent of the base datapath), seen once in 2^20; 16/75 at
3.22e-3. The GPU's energy moves the same way (its IMAD is the chain's own op), so the chip-against-GPU gain of every
weight corner is 1.0x, with 0 memory effect. Renormalised sums were 75 on every draw (0 failures). Fold rotations: a triple
all equal 2.13e-3, both triples all equal 1.91e-6, all six equal 0; uniform over 1..31, rotation 0 never drawn; a wire
mux, 1.0x.
## (c) The 64-bit seeding of the day-key stream (hash lane)
`attack-f7 days` over days 0..131,072 (`census-2p20.log`). The day key `K` is `seed_words_from_bytes("igneum-day/" ||
day_le64)`: a calendar function, no chain state. All 256 bits of `K` enter the cache fill (spec 1.8.3, `K[0..7]` in every
block input), so the dataset depends on the full key; the mixer-constant stream (ROT, MUL, RC) is seeded from `K[0] |
(K[1] << 32)`, 64 bits, which is the spec's stated intent (spec 1.8.4).
| Quantity | Value |
|---|---|
| Days the chain can have | about 65,745 in 180 years at 1 block/s (2^16.0) |
| Distinct 256-bit keys K over 2^17 days | 131,072 (all) |
| Distinct 64-bit stream seeds over 2^17 days | 131,072 (0 duplicates) |
| Distinct (ROT, MUL, RC) tuples over 2^17 days | 131,072 |
| Birthday bound on a 64-bit collision among 2^16 days | 2^(32 - 65) = 2^-33 |
The spec intends 64 bits for the mixer-constant draw, and the truncation is not a reduction of the draw space the public report
would flag: at most 2^16 days are ever drawn, each a distinct calendar day with a distinct 64-bit seed (0 collisions in
2^17), so no two days share a mixer. One observation, within spec intent and recorded for the written argument of
`funding.md` B5 rank 6: the mixer-constant stream has 64 bits of seed entropy, so at most 2^64 distinct daily mixers are
reachable (not the ~2^1,047 nominal); this is not exploitable (the days used are 2^16, all distinct) and whether any
reachable tuple is weak is the separate weak-day census of row F4.
## Consequences per tier
The era draw and the day-key seeding are protocol-wide and do not differ by card tier: the load width and load count are
pinned, so every era is equally memory-bound and no 8, 12, 16 or 24/32 GB card is advantaged or disadvantaged by any draw
(the measured six-era hash-rate spread is 1.3 percent on the RTX 5090, 3.2 on the RX 9070 XT, 0.8 on the M5 Max,
`algorithm.md` 5.4). No drawn era parameter or day key makes a chip cheaper against a GPU: the richest datapath corner is
1.0034x and is shared with the GPU. The one operational consequence is for the protocol, not a miner tier: the era draw's
grinding resistance is not yet demonstrable because the 1-hour VDF and its certified checkpoint are not in the node, so
the freeze of the draw procedure and the C_era cut rule must wait on the VDF landing and the finality lane's delay-
soundness measurement.
## Gate line
- (a) harness: INCOMPLETE. No re-roll with a one-block delay (known-fail 0 of 6); a re-roll with no delay (known-pass 1 of
6). The real gate (no re-roll inside the 2 s window) rests on the 1-hour VDF, which is not in the node; written argument
above.
- (b) census: PASS. No era class with gain over 1.1x at any fraction (richest 1.0034x, M = 1, absent in 2^24); the draw is
a bijection on every sample and uniform in R, pos and the M bits.
- (c) 64-bit seeding: PASS within spec intent. The spec intends 64 bits for the mixer stream; 2^16 days are all distinct;
the one observation (2^64 reachable mixers) is recorded, not a flaw.
What a failure moves (plan 4.2 F7): the draw procedure or the C_era cut rule; a redraw rule for the era stream. Nothing in
(b) or (c) moves them. (a) moves nothing in shipped code but gates the freeze of the draw procedure on the VDF.

View file

@ -0,0 +1,259 @@
# F8. Uniformity censuses of the class v4 derivation
Attack-pass row F8 (`docs/plans/cryptanalysis.md` section 4.2; the pass record `docs/analysis/attack-pass-2026-10.md`).
Run 7 October 2026, 08:13 to 09:3x UTC (09:13 to 10:3x UK) on igneum-build-1. Verdict: **FINDING** (AP-F8-1 below).
The line-index census is a PASS at its full sample size; the cross-hash item histogram is not uniform, and the cause
is in the base program, inside the acceptance rule's blind spot.
## 1. Target
| Item | Value |
|---|---|
| Commit | `igneum-pow` at 924288d1 (`attack-pass`); the box built HEAD b2a411d1, whose `igneum-pow` is byte-identical (`git diff --stat 924288d1 HEAD -- igneum-pow` is empty) |
| Class | `--program-class v4`: `V4_CLASS` = `mx8+sh256x27`, generator 4, mixer x8, the era layout drawn inside the class, the shadow block of 256 instructions x 27 reps |
| Line index | `proto-metal/MEMHARD.md` section 1.6: `a = s[0] AND 0x003fffff`, 4,194,304 lines of 64 B, 8 dependent reads per item (`memhard.rs` `derive_items_mask`, `cache.line_const(s[0])`) |
| Item index | `verify.rs` `load_index`: `y = rotl(x * M, R)`, the site's window `(y & (MASK >> k)) \| off`, then `Layout::split` removes the four interleave bits; 2^24 items at the 2^28-word dataset |
| Reads per hash | 128 loads (16 sites x 8 iterations), so up to 1,024 cache lines per hash and 32,768 per warp. The spec's analytic bound of 832 lines per hash (`docs/spec/01-lottery-hash.md` line 347, 104 loads x 8) predates generator 2's fixed 16 load slots; the current bound is 128 x 8 = 1,024 |
| Prior figures | `chip-model-v3.md` section 1: "median 128.00 distinct" items per hash (the 20,000-program census); `weak-program-census-2026-10-03.md` line 291: 127.7 distinct addresses per hash under the proposed generator |
| Day | the devnet pack's day, `bind::day_bytes(20730)` (2026-10-04), day 0 of the growth schedule: a 2^26-word cache, a 2^28-word dataset. Census 1 uses days 20730 to 20745 |
| Programs | p1 = the devnet epoch-0 derivation (epoch seed and era seed both the genesis hash `edc4fa84...fb07`, program id `c120d7963abdcd96`, attempt 0); p2 and p3 = chain-shaped seeds from tag strings (section 4), attempts 1 and 0 |
## 2. Method and harness
Harness: `tools/attack/f8-uniform/` (crate `attack-f8`, a path dependency on `igneum-pow`, nothing in the library
modified). Built on the box through `tools/build-remote.sh`: sha256 `590668...f913` for the firings of 2.1,
`ef8042...11b2` for sections 3, 4.1 and the first item runs (flat null, `log/c-*`, `log/d-*`), `890955...fbd3` for the
window-model runs and the seed census (`log/e-*`). The committed source carries one later label fix (the
"uniform-on-window" entropy reference in the per-site line is 16 - k_off bits; the 09:04 UTC logs print 16 - 2 k_off). Box scratch `/srv/builds/igneum-wt-attack/target-attack-f8/`
(the name `target-*` is what the box's checkout clean spared at the time; the fix at b92a5fd4 now also spares
`attack-*`). Every run: `nice -n 10 taskset -c 22-27,70-75`, 12 threads, under `flock -s /srv/builds/_locks/measure`
in chunks under 3 minutes each (the longest, phase D, under 25 minutes).
Two mirrors, each trusted only while it agrees with the library bit for bit:
| Mirror | What it records | Agreement check | Result |
|---|---|---|---|
| `derive_traced`: `memhard::derive_items_mask` instruction for instruction (`mixer`, `round_key_mult`, `cache.line` from the library), the line index of every round kept | 8 line indices per item | every item of every day also derived by the library's `derive_items` and compared on all 16 words | 0 mismatches on 268,435,456 items (section 3) and on 16,777,216 items per table build (section 4) |
| `Mirror::warp`: `verify::interpret_warp_init` for the class v4 op set, dataset words from a table of the day's 2^24 items, the item index and the source register of every load kept | 128 item indices per lane, the source value's saturation per position | the 32 hashes of warp 0 to 63 and of every 997th warp compared with `Epoch::hash_warp` | 0 mismatches on 95 warps per program (section 4) |
Three censuses:
1. `lines`: all 2^24 items of each of 16 consecutive day keys (2^28 item derivations, 2^31 line reads), the full 2^22-line
histogram per round and pooled, the 2^16-bucket histogram (64 lines, one chained segment per bucket), a uniform
SplitMix64 control of the same size.
2. `warps`: 10^6 nonces (31,250 warps) of each of three programs: distinct lines and items per hash and per warp, the
cross-hash item histogram, per-position diagnostics, an attribution pass from the hottest items back to the load
positions that read them.
3. `warps` at 2^26 nonces on p1: the one-epoch cross-hash item histogram at 512 expected reads per item.
The tests, defined before the runs:
- **6-sigma test**: the largest (and smallest) bucket of a histogram within 6 sigma of its expectation, sigma =
sqrt(expectation). The gate's bucket is the 64-line segment for lines and the 64-item bucket for items. The
full-resolution histograms are reported beside a uniform control of the same size, because at a small mean the
Poisson tail puts the maximum of 4 million bins above 6 sigma by chance (control at mean 8: +6.72 sigma; at mean
32: +5.83; at mean 512: +5.61).
- **Hot-set test** (F8's definition, written for F9's reuse): sort items by read count; S_f = the share of all reads
on the top-f fraction of items, for f in {0.1%, 0.5%, 1%}; E_f = the same share on a control of the same size drawn
from the design's own null (flat uniform for lines; the window-weighted null for items, section 4.2); the excess
X_f = S_f - E_f. **A hot set exists at f when X_f >= f**: after the chance excess is removed, the top f of items
capture at least one extra proportional share, which is what an on-die copy of f of the items would have to win to
matter. X_f / f is printed as the gain in proportional shares. The acceptance-style form of the same metric (for
rule (c)'s 2,048 evaluations): per load position, the largest count of one masked address, and the count of
saturated (0 or 2^32 - 1) source values.
### 2.1 The harness fires (known-fail and known-pass)
| Plant | What it does | 6-sigma test | Hot-set test | Log |
|---|---|---|---|---|
| `quarter-lines` | line index masked to a quarter of its range | buckets64 largest +75.97 sigma (2,231 at mean 512), smallest -22.63: FLAGGED | X_1% = +3.54% (S 5.60% vs control 2.06%), X/f = 3.5 at every f: FLAGGED | `log/a1-lines-quarter.log` |
| `half-lines` | line index masked to a half | buckets64 largest +29.26 sigma: FLAGGED | X_1% = +1.25%, X/f = 1.25: FLAGGED | `log/a2-lines-half.log` |
| `const-item` | one constant item at the first load site (1/16 of reads) | items buckets64 largest +92,682 sigma: FLAGGED | X_0.1% = +6.38%, X/f = 63.8: FLAGGED | `log/a4-warps-const-item.log` |
| none, 2^22 items, one day | the real derivation at a small size | buckets64 largest +4.42 sigma, smallest -4.51: within 6 sigma (control +4.33) | X_f = -0.0004%, -0.0006%, -0.0010%: clear | `log/a3-lines-pass-small.log` |
Both tests fire on every plant and neither fires on the real line derivation. Log paths are under
`/srv/builds/igneum-wt-attack/target-attack-f8/`.
## 3. Census 1: the line index over 2^28 derivations (PASS)
Sample reached: 16 days x 2^24 items = 268,435,456 item derivations, 2,147,483,648 line reads into 4,194,304 lines
(512 expected per line, 32,768 per 64-line segment). Mirror mismatches against `derive_items`: 0 of 268,435,456.
Log: `log/b-lines-16days.log`; histograms `out/lines-d20730-n16-i24-none-buckets64.txt` (65,536 rows) and
`out/lines-d20730-n16-i24-none-full.u32le` (4,194,304 x u32).
| Histogram | Bins | Expected | Largest | Sigma | Smallest | Sigma | chi2/dof | Top 1% share |
|---|---|---|---|---|---|---|---|---|
| Pooled, 64-line buckets (the gate) | 65,536 | 32,768 | 33,645 | +4.84 | 31,998 | -4.25 | 1.00226 | 1.01427% |
| Pooled, full 2^22 lines | 4,194,304 | 512 | 639 | +5.61 | 402 | -4.86 | 0.99937 | 1.11979% |
| Control, 64-line buckets | 65,536 | 32,768 | 33,524 | +4.18 | 31,960 | -4.46 | 0.99930 | 1.01426% |
| Control, full 2^22 lines | 4,194,304 | 512 | 639 | +5.61 | 408 | -4.60 | 0.99952 | 1.11956% |
| Per round 0 to 7, full, pooled (64 per line) | 4,194,304 | 64 | 107 to 113 | +5.38 to +6.12 | 26 to 29 | -4.75 to -4.38 | 0.99855 to 1.00062 | 1.3483% to 1.3488% |
| One day (20730), 64-line buckets | 65,536 | 2,048 | 2,291 | +5.37 | 1,859 | -4.18 | 0.99985 | 1.05826% |
| One day, control, 64-line buckets | 65,536 | 2,048 | 2,250 | +4.46 | 1,874 | -3.84 | 1.00377 | 1.05901% |
Per day, the gate bucket's largest value ran +4.00 to +5.37 sigma on all 16 days (control +4.46), every day within
6 sigma. Hot-set test on the pooled lines: X_0.1% = -0.00002%, X_0.5% = +0.00010%, X_1% = +0.00023% (X/f under
0.0003): clear. Round 0, whose input is the sequential item index through the init `t * MUL[i] + RC[i]` and eight
mixer applications, is as flat as rounds 1 to 7 (chi2/dof 0.99926; its +5.50 sigma maximum is below the control's
+5.61 at the pooled size). Round 5's +6.12 sigma at mean 64 is one bin of 4 million at a Poisson tail where the
control at mean 8 reached +6.72; its chi2/dof is 0.99855.
Gate line: the largest bucket is within 6 sigma of uniform (+4.84 on the 64-line buckets, +5.61 on the full 2^22
lines, both at or below the control), chi2/dof 0.99937, no hot set. **PASS at 2^28 derivations.**
## 4. Census 2 and 3: distinct lines per hash and warp, and the cross-hash item histogram
Setup per program: the day's 16,777,216 items derived once into a table with their 8 lines (6 to 9 s on 12 threads,
0 mismatches against `derive_items` on every item), then the warps interpreted from the table at 2.7 to 3.2 ms per
warp per thread. Logs: `log/c-warps-p{1,2,3}-1e6.log` (first run, flat null) and `log/e-warps-p{1,2,3}-1e6.log`
(windowed null, section 4.2); distributions `out/warps-<program>-d20730-n1000000-none-distinct.txt`, item histograms
`...-items.u32le` (16,777,216 x u32), per-position tables `...-positions.txt`.
### 4.1 Distinct lines and items per hash and per warp (10^6 nonces each)
| Program | Epoch seed / era seed | Lines per hash min / p1 / median / max / mean | Items per hash min / median / mean | Lines per warp min / median / max / mean | Items per warp min / median / mean |
|---|---|---|---|---|---|
| p1 `c120d7963abdcd96` (devnet epoch 0) | genesis / genesis | 1,008 / 1,023 / 1,024 / 1,024 / 1,023.867 | 126 / 128 / 127.9989 | 32,579 / 32,636 / 32,680 / 32,635.84 | 4,090 / 4,096 / 4,095.41 |
| p2 `82f0696f823e9c65` | `59cef1aa...bfdfa` / `9cba001f...1f69` | 1,014 / 1,023 / 1,024 / 1,024 / 1,023.871 | 127 / 128 / 127.9995 | 32,580 / 32,637 / 32,684 / 32,636.08 | 4,091 / 4,096 / 4,095.48 |
| p3 `e282eed7d47e425e` | `c54e2ddd...c95d` / `1b04f607...b58a` | 999 / 1,016 / 1,024 / 1,024 / 1,023.600 | 125 / 128 / 127.9656 | 32,276 / 32,481 / 32,601 / 32,480.32 | 4,050 / 4,076 / 4,075.85 |
| Uniform expectation | | 1,023.875 of 1,024 | 127.9995 of 128 | 32,640.3 of 32,768 | 4,095.50 of 4,096 |
Per hash, every program reads its 128 items and 1,024 lines as the design intends (p1 and p2 at the uniform
expectation; p3 a shade under, 127.97 items, which is the same site-15 effect as the finding below: the saturated
site repeats an item inside a hash 3 times in 100). Per warp, 32 lanes read 32,636 distinct lines of 2^22, a 2 MiB
working set of cache lines and 256 KiB of dataset items, within 0.01% of uniform on p1 and p2.
### 4.2 The cross-hash item histogram and the window layer
The era layout's window layer (`docs/plans/era-layout.md` section 1.4, layer 8) makes each load site read an
aligned half or quarter of the dataset with probability 2/3. The per-site item distribution is therefore not flat by
design (the diagnostic's "worst bit" reads P(1) = 1.0000 or 0.0000 at every windowed site: the fixed top bits), and
the summed item histogram has density steps between quarters. For p1 the 16 windows (site:shrink:offset
`7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1`) give
expected reads per item by quarter of 3.25 : 2.25 : 5.75 : 4.75 in sixteenths of the flat value. Against a flat
uniform the 64-item buckets of p2 (a program without the finding) read +10.03 and -9.06 sigma, which is the window
layer and not a flaw. The item tests are therefore judged against the **window-weighted null**: the expected count of
every item from the program's 16 windows, and a control that draws each read from a uniformly chosen site's window.
A chip gains nothing from the window steps: the union of the windows is the whole dataset every hour (era-layout.md
section 7), the floor window is 2^26 words (256 MiB), and which quarter is dense changes with the program.
#### The window model (reproducible by any reader)
For load site s with window draw `(k_s, o_s)` at the 2^28-word dataset: `k = min(k_s, 28 - 26)`, the word window is
`[o_s << (28 - k), (o_s + 1) << (28 - k))`; the item window is `[o_s << (24 - k), (o_s + 1) << (24 - k))` of
`2^(24 - k)` items (the four interleave positions all lie below bit 16, so the top bits of the word index are the top
bits of the item index). The expected reads per item is `E[t] = sum over sites s with t in window_s of N x 8 / 2^(24 - k_s)`
for N nonces (8 iterations per site), a density constant on each quarter of the item space. The windowed control draws
each of the N x 128 reads as (site = read index mod 16, item uniform on that site's window). Both controls are drawn from
SplitMix64 with a fixed seed. The tests on items are run against E[t] (chi-square, sigma of the largest and smallest
64-item bucket) and against the windowed control (the top-f shares); the flat uniform numbers are kept beside them as
what an auditor sees first.
#### Results, 10^6 nonces per program, 128,000,000 reads (`log/e-warps-p{1,2,3}-1e6.log`)
| Program | Windows (k_off:offset per site) | Quarter densities (reads per item) | Buckets64 largest sigma, windowed (control) | chi2/dof windowed (control) | Top 0.1% share: real / window control / flat control | Ratio to window control at 0.1% (gate 1.2x) | Ratio to flat control | Hot set (X_f >= f) |
|---|---|---|---|---|---|---|---|---|
| p1 devnet epoch 0 | 2:1 1:1 1:1 1:1 0 1:1 0 2:2 1:1 1:1 2:0 0 0 0 2:0 1:1 | 6.20 / 4.29 / 10.97 / 9.06 | +45.77 (+4.95) | 1.2336 (0.9981) | 0.5458% / 0.2891% / 0.2429% | 1.888x BEYOND | 2.247x | yes at 0.1% (X/f 2.57) and 0.5% (1.20); not at 1% (0.46) |
| p2 | 2:0 1:0 0 0 2:3 2:2 1:0 0 2:3 0 0 0 0 1:1 2:0 0 | 9.54 / 5.72 / 6.68 / 8.58 | +4.59 (+4.64) | 1.0061 (0.9986) | 0.2716% / 0.2639% / 0.2429% | 1.029x within | 1.118x | no (X/f 0.08, 0.05, 0.05) |
| p3 | 1:0 0 0 2:2 0 0 1:0 1:0 1:0 2:2 2:2 1:1 0 0 0 0 | 7.63 / 7.63 / 10.49 / 4.77 | +12,245.66 (+5.06) | 907.67 (0.9993) | 4.5954% / 0.2792% / 0.2433% | 16.46x BEYOND | 18.92x | yes at every f (X/f 43.2, 9.9, 5.0) |
p2 is what the class is designed to be: against the window model its largest bucket is +4.59 sigma (the control +4.64),
chi2/dof 1.006, the top 0.1% of items hold 1.029x their window-model share, and the flat-control ratio of 1.118x is
the window layer. p1 and p3 are the finding (section 5). The one-epoch histogram at 2^26 nonces (8,589,934,592 reads,
512 per item, `log/d-warps-p1-2e26.log`, flat null): p1's top 0.1% hold 0.5199% of reads against 0.1152% flat
control (X/f 4.05), the top 1% 2.4946% against 1.1198% (X/f 1.37), item 0xca5b92 78,479 reads at a mean of 512, and
site 15 feeds 6.37% of its reads into the top 0.1% in each of the 8 iterations; the excess grows with N as the
control's chance excess shrinks, which is the signature of a structural skew. Distinct lines and items per hash and
per warp at 2^26 nonces: 1,023.866 / 127.9989 / 32,635.6 / 4,095.41, unchanged from 10^6.
## 5. AP-F8-1: a saturated load source makes a cross-hash hot set (FINDING)
**What**: an accepted class v4 program can read one load site from a register whose last writes after its last
injecting write are `or` (and, mildly, `mul`), so the site's address has fewer than 32 bits of entropy across nonces
and the same items are read by many hashes. The per-hash figures (128 distinct items, 1,024 lines) stay intact; the
cross-hash item histogram does not. It is not the window layer (p2 shows the window layer alone is clean against its
model) and not the shadow block (iteration 0's load, which runs before any shadow block, is as hot as iterations 1 to
7: p1 6.372% vs 6.371% to 6.378%; p3 71.9% vs 72.4% to 72.6%).
**Where it hides from rule (c)** (`accept.rs`, 2,048 evaluations of the base program): the tests are constant bits
in FINAL register values, one address in ALL 32 lanes of a unit, saturated FINAL values, output-bit bias, and distinct
addresses WITHIN a hash. A site whose address is concentrated across hashes but refreshed before the end of the
iteration passes every one. Rule (a) accepts any write, `or` included, as the refresh between two loads from the same
register (`check_stale_loads`); `Op::injects` (add, sub, xor, mad, shfl, load) is only used by rule (b), once per
register per program.
**The index derivation at the hot site** (the "writers back to the last injecting one" lines of `log/e-warps-p*.log`):
| Program | Hot site | Source | Writes after the last injecting write | Site's reads into the top 0.1% of items (flat expectation) | Index entropy, 256-item buckets (uniform on window) | Saturated source (x = 0 or 2^32 - 1) | Most repeated address at one position in 2,048 evaluations (uniform: 1 to 2) |
|---|---|---|---|---|---|---|---|
| p3 | site 15, instr 62 | r5 | `load@17` then `or@19`, `or@30` | 72.43% (0.10%) | 13.411 bits (16) | 1.368% | 44 of 2,048; 32 saturated |
| p1 | site 15, instr 63 | r6 | `add@51` then `rotl@53`, `or@61` | 6.93% (0.11%) | 14.985 bits (15) | 0.005% | 2 of 2,048; 0 saturated |
| p2 (clean) | every site | | injecting, or bijective (`rotl`), or `mul`/`mulhi` | 0.41% to 0.97% (0.20%; the window densities) | 13.999 / 14.997 / 15.994 bits (14 / 15 / 16) | 0.000% | 2 of 2,048; 0 saturated |
In p3 two `or`s on r5 after its load make the source 1 with probability 7/8 per bit; x = 2^32 - 1 in 1.37% of
evaluations and the images of the near-saturated values under the stride (`y = rotl(x * M, R)`, 256 x-values per
item) pile onto a few items: 0xffdf69 takes 213,913 of the site's 8,000,000 reads (2.67%), the top 0.1% of items
72.4%, and 4.6% of ALL reads of the hash land on 0.1% of the items. In p1 one `or` after `rotl(add)` gives 3/4 per bit
on the ORed positions: no saturation to speak of (0.005%), but 6.9% of the site's reads on 0.11% of the items (the
hot items share the low 20 bits `5b92`: 0xca5b92, 0x8a5b92, 0xaa5b92, 0xba5b92, 0x825b92, 0xe65b92, 0x985b92), a
2.6x proportional excess at f = 0.1%. p2's `mul` sites (10, 13: `mul` after a load or a shuffle) read 0.65% and 0.70%
into the top 0.1% against 0.41% and 0.48% for their window class (an even multiplier zeroes low bits; the hot items
0xd6a680, 0xe44400, 0xd25600 end in zero bits), a mild effect that the window-model ratio (1.029x) absorbs.
**How common** (the seed census, 64 chain-shaped programs p4 to p67, 262,144 nonces each, `log/e-seed-census-4-67.log`,
`out/seed-census-d20730-n262144-p4-67.txt`): CENSUS-LINE
**Reproduction**: `attack-f8 warps --program 3 --nonces 1000000 --diag 1` (or `--program 1`); the acceptance-style
numbers come from the same run's "acceptance-style" line. The program is `Epoch::chain_program(epoch_seed, Some(era),
ProgramClass::V4, label)` with the seeds of section 4.1.
**Proposed fix** (not applied; `igneum-pow` untouched, the Counter ASIC lane re-gates on `ca3-v4-uniform` with this
harness):
1. Rule (a'), static: between the last injecting write of a load's source register and the load (cyclically), no
`or` and no `mul` writes that register; `rotl`, `rotr` and `mulhi` may (bijective, or measured flat: p2 site 15
reads `mulhi` after `add` at 1.00x). This rejects p1 and p3 at draw time and costs nothing at run time. Programs
rejected are redrawn as today (`MAX_ATTEMPTS` 32); the census gives the rejection rate.
2. Rule (c'), dynamic, the same 2,048 evaluations: no load site reads a saturated source (0 or 2^32 - 1) in more
than 2 evaluations, and no address repeats more than 4 times at one position (uniform expectation 1 to 2; p3 shows
44 and 32). This catches the strong class only; p1's class needs about 2^16 evaluations to show at a site (65,536
nonces: largest item count 67 at a mean of 0.5), so (a') is the rule that closes it and (c') is the check that
fails loudly if (a') is ever loosened.
3. Packs re-cut for the seeds the new rule rejects (the devnet epoch-0 program p1 is one of them: its site 15 is
`or@61`), with the gate pack ids re-pinned; the chain's own epochs redraw automatically.
**Reuse for F9**: the hot-set metric (section 2) on the per-program item histogram at 2^18 nonces, and the
acceptance-style pair (most repeated address at a position, saturated sources at a position) at 2,048 evaluations,
are both emitted by `warps --programs a..b`; a header-grinding search that steers a program to a hot set would show as
ratio-to-window-model above 1.2x at f = 0.1%.
## 6. Consequences per tier
| Number | What it means | Per tier |
|---|---|---|
| Line index uniform at 2^28 derivations (largest segment +4.84 sigma, chi2/dof 0.99937) | the 256 MiB cache has no hot segment: a chip or a card cannot serve the 8 dependent reads of an item from a cache smaller than the whole 256 MiB (the floor window of era-layout.md) | no change for any card; the verifier's cache stays 256 MiB in RAM on every node |
| Distinct lines per hash 1,023.87 of 1,024, items 127.999 of 128 (p1, p2); per warp 32,636 lines, 4,095 items | the per-hash working set is 64 KiB of cache lines and 8 KiB of items, per warp 2 MiB of lines and 256 KiB of items; the item-derivation chip's "128 items per hash" input (`chip-model-v3.md`) stands | the 8 GB card and up: unchanged; the recompute chip pays 128 derivations per hash, as modelled |
| p3-class programs: 4.6% of all dataset reads on 0.1% of items (1 MiB of a 1 GiB dataset); p1-class: 0.59% on 0.11% | a stored-dataset chip with 1 MiB of on-die SRAM serves 4.6% of its reads without touching DRAM on such an epoch; a GPU's L2 (96 MiB on the 5090, 64 MB Infinity Cache on the 9070 XT, vendor figures) holds the same 1 MiB, so both sides gain the same 4.6% of reads and the chip's edge from it is about 0 (the per-joule edge of `evidence.md` row 17 is a DRAM-read figure; a 4.6% read saving on both sides moves it by under 5% on such epochs). The recompute chip (f = 0, SRAM cache) caches the derived hot items and skips up to 4.6% of its 128 derivations per hash on such epochs, a 4.8% rate gain on those epochs only | home cards 8 to 32 GB, rigs, pools: no action; a few percent of epochs run a few percent faster for everyone with an L2. The verifier: `MemhardCpu::fetch` dedupes within a fetch only, so no change. The chip model: the headline 2.1x at k = 1 moves by under 5% on affected epochs and 0 on others; the fix below returns it to 0 everywhere |
| The acceptance rule's blind spot (cross-hash concentration at one site) | a program class property, not a day or era property: the same seed is hot on every day and under every era, so a chip or a pool that selects epochs cannot gain more than the epoch's own 4.6%; but the public claim "the item map is uniform per program up to the window layer" is false for the affected fraction of seeds until rule (a') lands | the fix is a generator rule plus packs re-cut: a class change under the 95% signalling rule if it lands after the flip, a plain re-cut if it lands in the class v4 cut itself (the lane's call) |
## 7. Gate line and verdict
| Gate (plan 4.2 F8, the same as 1.4 (4)) | Result | Status |
|---|---|---|
| The largest bucket within 6 sigma of uniform on the stated sample sizes (line index, 2^28 derivations) | +4.84 sigma on 64-line segments, +5.61 on 2^22 lines (control +4.18 / +5.61), chi2/dof 0.99937 | PASS |
| The item distribution within 6 sigma of uniform (against the window model, the design's own null) | p2 +4.59 sigma (control +4.64); p1 +45.77; p3 +12,245.66 | FAIL on p1 and p3 |
| No hot set under 1% of items among passing seeds (10^6 nonces on three programs; the 64-seed census at 2^18) | p2 none; p1 top 0.1% at 1.888x the window model (2.247x flat), X/f 2.57; p3 16.46x (18.92x flat), X/f 43.2; census: 31 of 64 seeds over 1.2x of the window model, 23 with a hot item | FAIL |
| The Counter ASIC lane's record gate: top 0.1% within 1.2x of the window-model control on every seed | p2 1.029x; p1 1.888x; p3 16.46x; census: 31 of 64 seeds over 1.2x (median 1.17x, p90 2.70x, max 13.09x on p31) | FAIL |
**Verdict: FINDING (AP-F8-1).** The line index passes at 2^28 derivations. The cross-hash item histogram fails
the hot-set gate on 2 of the 3 named programs (one of them the live devnet epoch-0 program) and on 31 of 64 seeds (48 percent) of of
the 64-seed census, from `or` (and mildly `mul`) writes on a load's source register after its last injecting write,
outside every test of rule (c). What it moves: not the mask or the fold (the derivation is uniform) but the
acceptance rule, (a') and (c') above, and the packs re-cut. Ownership: the Counter ASIC lane (generator and rule),
re-gated with this harness on the fixed branch; the row reads FIXED-AND-PASSED when every seed of the census passes
both the hot-set test and the 1.2x gate under the new rule.
Sample sizes reached: 2^28 derivations (lines); 10^6 nonces on three programs (distinct lines, hot set); one epoch at
2^26 nonces (cross-hash histogram); 64 seeds at 2^18 nonces (the census).
Times UTC in the logs; the runs ran 08:13 to 09:2x UTC on 7 October 2026 (09:13 to 10:2x UK).

View file

@ -0,0 +1,301 @@
# F9: acceptance edges, the hot-set search, header grinding
Attack-pass row F9 of `docs/plans/cryptanalysis.md` section 4.2 (record: `docs/analysis/attack-pass-2026-10.md`).
Sub-agent attack-f9, 7 October 2026. Status: IN PROGRESS (rewritten as each run lands; the numbers below are the
ones already final, each with its log).
## Target
| Item | Value |
|---|---|
| Commit | 924288d1 (branch attack-pass, worktree igneum-wt-attack) |
| Generator | 4, class v4 `mx8+sh256x27` composed with the era draw (`LoadClass::era(V4_CLASS, E, [4 bytes])`), era seed E = the devnet epoch-0 seed `edc4fa84...fb07` (`proto-cuda/packs-ca3-v4/v4-devnet-epoch0/seeds.txt`) |
| Rule | `igneum-pow/src/accept.rs`: (a) stale load sources, (b) injecting writes, (c) the 2,048-evaluation dynamic test on the closed-form stand-in `dataset_elem` at 2^28 words with init words = seed words; redraw on rejection up to 32 attempts |
| Header binding | `igneum-pow/src/bind.rs`: init words `I = seed_words_from_bytes("igneum-block/" \|\| H \|\| nonce_hi_le32)`, one `I` per 32-lane warp, the lane nonce in the low 32 bits |
| Memory-hard dataset for the edges | the devnet day 20730 (`day_seed_hex 69676e65756d2d6461792ffa50000000000000`), class v4 shape (mixer x8, cache 2^26 words, dataset 2^28 words), `Epoch::chain_dataset_day` |
| Card | RunPod RTX 5090 (170 SMs, 32,120 MiB, driver 570.195.03, CUDA 12.8.1), pack `v4-devnet-epoch0` built there with `nvcc -O3 -arch=sm_120` |
## Known-failed shape
A seed grind that steers a program to a hot cache set for DRAM locality, or an edge where the closed-form stand-in
disagrees with the live verifier in the attacker's favour.
## Gate
Zero passing programs with a hot set under 1 percent of items among 10^6 seeds; the grinding gain under 1 percent of
rate at any search cost. What a failure moves: the closed-form stand-in replaced by the live verdict at the edges; a
locality term in rule (c).
## Harness
`tools/attack/f9-grind/` (crate `attack-f9`, `igneum-pow` as a path dependency, nothing in igneum-pow edited; built
on igneum-build-1 through `tools/build-remote.sh`; the binary on the box at
`/srv/builds/igneum-wt-attack/tools/attack/f9-grind/target/release/attack-f9`):
| Sub-command | What it does |
|---|---|
| `selftest` | the firings listed below |
| `edges` | sub-row (a): every candidate of every seed through the re-implemented dynamic test twice, closed form and memory-hard, every metric of rule (c) measured to the end (no early abort) with its margin; the attempts continue until both stand-ins have accepted, so the chosen program under each is known |
| `hotset` | sub-row (b): the seed's accepted program, its 2,048 x 128 address record under the acceptance init and under a block init; per site the nonce-independent address bits, the distinct addresses and the most-read address; the histogram at bucket scales 2^8 to 2^24 words against a window-aware Poisson expectation (the era windows send a site to the dataset, a half or a quarter of it) with a Bonferroni tail; the taint count of init-determined loads |
| `inspect` | one seed's program with every site that repeats an address |
| `grind`, `table-random` | sub-row (c), CPU side: the init-determined load sites of the devnet epoch-0 program, the per-warp search over K nonce_hi values for the fewest distinct 128 B lines inside those load instructions (`--mode intra`, what the coalescer merges) or across them (`lines`, `pages`), the search cost per hit, and the per-warp init tables the card reads |
| `reference` | the 64 bound hashes the card's known-pass compares against; the file used on the pod came from the pre-built `igneum-pow hash-bound` instead (`ref.txt`, sha256 e831458a...) |
| `summarise.py` | the census summaries quoted below |
`tools/attack/f9-grind/pod/` (the card): `make-variants.py` copies the pack's `kernel_bound.cu` into five kernels
(per-warp init table; the five init-determined loads broadcast to lane 0's address; every load broadcast; the first
such load broadcast; lane 1 reading lane 0's address at the first such load), `f9-host.cu` fills the cache and dataset
with the pack's own kernels, checks them against `vectors.h`, checks the bound hash against the reference, checks the
per-warp kernel on an all-equal table against the honest kernel, then times the eight variants in interleaved rounds;
`run.sh` builds on the pod, samples `nvidia-smi` once a second and joins the samples to the phases (`join-power.py`).
Hot-set metric (F8's record `docs/analysis/attack-pass/f8-uniform.md` did not exist when this harness was written, so
the metric is defined here). Strict reading, "any hot bucket": a site with 7 or more nonce-independent address bits
(support at most 2^21 of 2^28 words, 0.78 percent of items; the window's own fixed bits not counted), or any bucket
of at most 2^20 words (0.39 percent of the dataset) at scales 2^8, 2^12, 2^16, 2^20 whose count has a Poisson tail
against its window-aware expectation under 10^-6 after the Bonferroni correction. Gate reading, "flagged": the reads
above expectation in those hot buckets (the hot share, what a cache of the hot set saves at most) reach 1 percent of
the program's reads, or a site has 7 constant bits.
## Firings (the harness is trusted only after these)
Log: `/srv/builds/igneum-wt-attack/attack-f9/selftest.log` (copy in the Mac scratchpad `f9-box/selftest.log`).
| Check | Known-pass | Known-fail | Result |
|---|---|---|---|
| 1 | the re-implemented dynamic test against `accept::check` on 300 class v4 candidates: every verdict, the first failing condition, and distinct, saturated and bias of every accepted report equal | | PASS (300 candidates, 16 rejected by accept, all equal, 1.5 s) |
| 2 | both stand-ins forced equal (closed form twice) on 200 candidates | | PASS, 0 disagreements |
| 3 | | the memory-hard stand-in gives different words: distinct 262,117 against 262,106, bias 56 against 64 on one program | PASS (they differ) |
| 4 | 50 accepted programs, none with 7 constant bits (worst 0) | 50 plants (an accepted program rewritten to `xor a,a; add a,a,256; mulhi a,b` before a load from `a`, 48 of 50 still pass rule (c)) all flagged (support 256 words, site distinct 255 or 256) | PASS for the plant; 4 of the 50 clean programs have a hot bucket (sub-row (b): that is the finding, not a harness fault) |
| 5 | taint on the devnet epoch-0 program against the hand reading of `kernel_bound.cu`: loads 7, 8, 9, 10 read r7, r4, r2, r0 (no load before them); load 31 reads r5 = r5 x r4 from instruction 12, both untouched by any load; loads 11, 13, 29, 30 read r1, r6, r4, r3, each written by an earlier load or by `mad` from r7 after load 10 | | PASS: sites (0,7) (0,8) (0,9) (0,10) (0,31) |
| card 1 | cache FNV 448274a57f508cbc, dataset head, last word and 64 samples, the 96 pack vectors | | PASS (`pod log/host.log`) |
| card 2 | the bound hash against the 64 reference lines of `igneum-pow hash-bound` (nonce_hi 0, prehash 000102..1f) | | PASS, 0 wrong |
| card 3 | the per-warp kernel on an all-equal table equals the honest kernel on 64 lanes | the two-init table: warp 0 equal, warp 1 differs in 32 of 32 lanes | PASS both |
| card 4 | | the forced kernels change the hash: forced4 64 of 64 lanes, forcedall 64, forced1 64, pair 64 | PASS (they fire) |
| card 5 | | a deliberately locality-maximising choice shows a measurable change: `pair` (one line of 4,096 saved per warp) +0.21 percent, `forced1` (31 lines) +6.8 percent, `forced4` (155 lines) +43.3 percent, `forcedall` +181.9 percent in the smoke run | PASS (measurable from one saved line up) |
## Sub-row (a): the edges
Run: `attack-f9 edges` over seeds `igneum-f9/0` to `igneum-f9/99999`, 8 threads on cores 28-31,76-79 in 10,000-seed
chunks under a shared hold of the box measure lock (`run-census.sh`); output `edges.part*.tsv`, summary by
`summarise.py edges`. The first 20,000 seeds (parts 0 and 1) are summarised here; the full 10^5 replaces this table
when the run ends.
| Quantity | First 20,000 seeds |
|---|---|
| Candidates evaluated | 21,020 |
| Verdicts agreeing | 21,007 |
| Disagreements | 13 (0.062 percent of candidates; the 3 October census had 39 in 100,000 on its generator) |
| Seeds whose chosen program differs | 13 (every disagreement moves the chosen attempt, because the next attempt was accepted by both) |
| Exhausted seeds | 0 on either stand-in |
| Rejected by the closed form / by the memory-hard dataset | 1,013 / 1,014 (850 static, the rest (c)) |
| First failing (c) condition, closed form | const_bit 80, saturated 54, distinct 24, lane_const 4, bias 1 |
| Accepted margins, closed form | saturated at most 81 of 164, bias at most 120 of 136, distinct sum at least 247,335 (bound 245,760), nearly constant final bits up to 2,047 of 2,048 |
| Accepted margins, memory-hard | saturated at most 82, bias at most 115, distinct at least 247,678 |
The 13 disagreements: 12 are `const_bit`, a final register bit equal in all 2,048 evaluations on one dataset and in
2,044 to 2,047 of them on the other (7 where the closed form accepts, 5 where the memory-hard dataset accepts); 1 is
`bias`, output bias 155 against 93 (6.9 against 4.1 sigma, the two draws' difference 2.7 sigma of sampling noise),
where the memory-hard dataset accepts. No disagreement on saturation, lane-constant sites or the distinct count: those
metrics are the same to within 20 on both datasets. What an attacker gains from a program the closed form accepts
and the live dataset would reject: a register whose final bit is pinned in 2,047 of 2,048 hashes instead of 2,048,
which no test downstream of the fold can see (the 64 output bits stay within 120 of 1,024 on every accepted program)
and which no chip can turn into skipped work; the reverse direction loses the chain a program with one pinned bit.
Either way the chosen program moves to the next attempt, which both stand-ins accept. Nothing in the attacker's
favour: the verdict's dependence on the stand-in is a 0.06 percent coin flip on a one-bit property.
## Sub-row (b): the hot-set search
Run: `attack-f9 hotset` over seeds `igneum-f9/0` to `igneum-f9/999999`, 8 threads on cores 32-35,80-83 in
100,000-seed chunks; output `hotset.part*.tsv`. A 2,000-seed timing sample (`hotset-timing.tsv`, seeds 5,000,000 to
5,001,999) is summarised here; the 10^6 census replaces it when the run ends.
| Quantity | 2,000-seed sample |
|---|---|
| Programs with any hot bucket (strict) | 161 of 2,000 (8.1 percent) |
| Programs flagged at the gate reading (hot share at least 1 percent of reads) | 21 of 2,000 (1.05 percent) |
| Worst hot share | 10.3 percent of the program's reads (seed 5,000,968) |
| Sites with 7 or more constant address bits | 0 (max 0) |
| Fewest distinct addresses at a site in 2,048 evaluations | 434 |
| Most evaluations reading one address at a site | 767 of 2,048 |
| Init-determined loads in iteration 0 (programs by count) | 1: 234, 2: 573, 3: 594, 4: 368, 5: 179, 6: 42, 7: 10; none after iteration 0 |
FINDING F9-1 (or-saturation hot words). `inspect --seed 5000968` (`inspect-5000968.log`): the load at instruction 33
reads r4; r4 is written by `or r4 |= r0` (20), `mulhi` (27) and `or r4 |= r6` (28), and r6 itself by `or r6 |= r7`
(7). `or` is absorbing toward all ones: after two `or` writes from independent words every bit is set with
probability 7/8 and the whole register with probability (7/8)^32 = 1.4 percent; chained across iterations the mass
grows, and on this program r4 is 0xffffffff at that site in 731 to 739 of 2,048 evaluations in iterations 1 to 7
(36 percent). The site then reads one word, `rotl(0xffffffff x M, R) & window | offset` = 0x0ca59e4c for a full
window and 0x04a59e4c, 0x08a59e4c, ... for the windowed sites; near-all-ones values add a few hundred more. The
same word family appears in every flagged program (seeds 4,000,001, 4,000,037, 4,000,040 in the selftest: `or`
writes at 54 and 55 before the load at 60, or at 1 before the load at 3). Rule (c) does not see it: the saturation
test counts final register values only (the register is overwritten before the end), the lane-constant test needs
all 32 lanes equal, the distinct test counts per lane per hash (the hot word repeats across iterations, so it
costs one distinct of 128), and the output bias stays within tolerance. The 3 October census measured an
`or_sat_frac` per program (section 7.3, max 0.0102) but the adopted rule kept only the final-value count.
What it is worth to an attacker: nothing asymmetric. The hot words are the same for every lane that saturates, so
the GPU's coalescer and L1 already serve them without a DRAM transaction, and a chip gets exactly the same. What it
costs the design: those programs do fewer memory-hard reads than rule (c) promises (up to 10 percent fewer on the
worst program in 2,000, at least 1 percent fewer on about 1 program in 100), so the per-hash memory work of class
v4 is not the uniform 128 random reads the chip model assumes on every epoch. Gate reading: FAIL in the strict
reading (zero passing programs with a hot set under 1 percent of items), FAIL in the share reading too (programs
with a hot set capturing at least 1 percent of reads exist at about 1 percent of epochs). Proposed fix, for the hash
lane (not applied here): a per-site line in rule (c), "every load site reads at least 2,000 distinct addresses over
the 2,048 evaluations" (uniform gives 2,048 minus 0.008 expected repeats; the saturated sites read 434 to 1,855),
computed from the addresses the test already collects (one sort of 2,048 per site, 128 sites, under a millisecond);
the redraw rate rises by about the strict-reading fraction (8 percent of candidates) unless the threshold is placed
at the share reading. The alternative, dropping the `or` family from the draw table, changes the frozen weights and
is for the lane to weigh. Class check: a chip gains nothing today, but a stand-in that lets 1 percent of epochs run
with a 1 to 10 percent lighter memory side is a published-number problem (evidence row 17's per-hash reads).
(the 10^6 numbers and the hot-share distribution replace the sample when the census ends)
## Sub-row (c): header grinding
### What an attacker can steer
Only a load whose address register has not yet absorbed a dataset word is a function of the init words and the
nonce alone (taint analysis, `init_determined_sites`). On the devnet epoch-0 program these are the loads at
instructions 7, 8, 9, 10 and 31 of iteration 0; from iteration 1 every register is tainted. Over the 2,000-seed
sample the count is 1 to 7 per program, median 3, always in iteration 0 only. Everything after depends on dataset
words the miner must fetch first. The init words themselves are an FNV hash of the header and nonce_hi, so the
attacker cannot choose them, only draw them; and one draw serves a whole warp (the shuffles couple the 32 lanes),
so a per-lane draw costs 32 hashes per lane.
### The search (CPU)
`grind` draws K init words per warp (nonce_hi 0 to K-1 under the fixed prehash) and keeps the one with the fewest
distinct 128 B lines among the init-determined loads. Each try costs 32 lanes x (8 init + 32 prefix instructions) =
1,280 lane-instructions; the warp's hash costs 32 x (512 + 55,296) = 1,785,856 lane-instructions, the derivation
not counted. Logs: `grind-k10-crosssite.log` (lines counted across the five sites: coincidences that are at best an
L2 hit), `grind-k10-pages.log` (2 KB pages across the sites), and the `intra` runs (lines inside one load
instruction, what the coalescer merges into one transaction) that feed the card.
| Metric | K | Warps | Mean lines or pages saved per warp (of 4,096 loads) | Warps improved | Search per warp in hashes |
|---|---|---|---|---|---|
| lines across the sites | 2^10 | 524,288 | 0.895 (0.022 percent) | 89 percent | 0.73 |
| 2 KB pages across the sites | 2^10 | 65,536 | 1.455 (0.036 percent) | 98 percent | 0.73 |
| lines inside one instruction (intra) | 2^10 | 2^17 | (pending) | | 0.73 |
| lines inside one instruction (intra) | 2^14 | 2^17 | (pending) | | 11.7 |
### The card (RTX 5090)
Smoke run (1 round of 2 s per variant, `pod smoke` logs): the calibration of what one saved line is worth.
| Variant | What changes | MH/s | Against honest | W | MH/J against honest |
|---|---|---|---|---|---|
| honest | the pack kernel, one init per dispatch | 141.76 | | 482 | |
| perwarp-random | per-warp init table, no search | 141.73 | -0.02 percent | 486 | -0.7 percent |
| perwarp-k10 | per-warp table, best of 2^10 (cross-site table) | 141.74 | -0.01 percent | 487 | -1.0 percent |
| perwarp-k14 | (the same table in the smoke run) | 141.74 | -0.01 percent | 488 | -1.1 percent |
| pair | lane 1 reads lane 0's address at load 7: 1 line of 4,096 saved | 142.05 | +0.21 percent | 490 | -1.4 percent |
| forced1 | load 7 broadcast: 31 lines saved | 151.39 | +6.8 percent | 506 | +1.8 percent |
| forced4 | loads 7, 8, 9, 10, 31 broadcast: 155 lines saved | 203.11 | +43.3 percent | 530 | +30 percent |
| forcedall | every load broadcast: 3,968 lines saved | 399.61 | +182 percent | 520 | +161 percent |
Reading: the class v4 kernel on the 5090 is bound by its random reads (141.8 MH/s x 128 = 18.1 G reads per second,
the card's measured random-read ceiling in `docs/bench-log.md`), and a load instruction completes when its slowest
lane's transaction returns, so one saved line is worth about 0.2 percent of rate, 31 lines 6.8 percent, and the
five init-determined loads fully coalesced 43 percent. That ceiling is unreachable by search: it needs the 32
lanes' 28-bit addresses to fall in one line at five sites, probability 2^-115 per draw. What a draw can reach is one
coalesced pair at one site (probability 5 x C(32,2) / 2^23 = 3 x 10^-4 per try, about 3,400 tries per pair);
two pairs need about 6 million tries, m pairs about 3,400^m / m! tries. One pair is worth 0.2 percent of one warp's
hash and costs 3,400 x 1,280 lane-instructions = 2.4 hashes of search. The measured per-warp tables (5 rounds of
8 s, pending) are the direct check.
(the full run's table replaces the smoke run when it ends)
## Consequences per tier
(filled in with the verdict)
## Logs
| Log | Path |
|---|---|
| selftest, inspect, grind, census parts and drivers | `/srv/builds/igneum-wt-attack/attack-f9/` on igneum-build-1 (`selftest.log`, `inspect-*.log`, `grind-*.log`, `edges.part*.tsv`, `edges.driver.log`, `hotset.part*.tsv`, `hotset.driver.log`, `hotset-timing.tsv`, `ref.txt`, `table-*.bin`) |
| card smoke run and full run | the pod's `/workspace/f9/podjob/smoke/log/` and `log/` (`run.log`, `nvcc.log`, `host.log`, `power.csv`, `power-by-variant.txt`, `sha256.txt`), copied to `/srv/builds/igneum-wt-attack/attack-f9/pod/` at the end |
### The card, the measurement (RTX 5090 pod ap-f9, `pod/host.log`, sha256 83b1372c..., copied to
`/srv/builds/igneum-wt-attack/attack-f9/pod/`)
Per-warp header grinding at K = 2^14 draws per warp against the honest kernel, 5 interleaved rounds of 8 s each:
141.62 against 141.61 MH/s, +0.004 percent of rate, sd 0.003, at 11.7 hashes of search per hash. The unreachable
ceiling (the five init-determined loads fully coalesced, `forced1` extended) is +43 percent. Gate: the grinding gain
under 1 percent of rate at any search cost. Sub-row (c): PASS. Pod time about 1 h 50 min from 09:02 UK; destroyed on
the lane's done line.
## Full censuses (box 1 parts 0 to 7 and 0 to 6; box 2 parts `edges-b2`, `hotset-b2` after the lane's move to
build-2 at 12:2x UK; `summarise.py` over all parts, 12:5x UK)
### (a) The edges on generator 4, 10^5 seeds
| Quantity | 100,000 seeds |
|---|---|
| Candidates evaluated | 105,064 |
| Verdicts agreeing | 105,030 |
| Disagreements | 34 (0.032 percent of candidates; the 3 October census had 39 in 100,000 on its generator) |
| By kind | `const_bit` 29 (closed form accepts 17, memory-hard accepts 12); `bias` 4 (3 and 1); `lane_const` 1 (memory-hard accepts) |
| Seeds whose chosen attempt differs | 34, every one moving to the next attempt, which both stand-ins accept; exhausted 0 on either |
| Rejected, closed form / memory-hard | 5,046 / 5,048 (static 4,201; then const_bit 424 / 428, saturated 275 / 275, distinct 112 / 112) |
| Accepted margins, closed form | saturated at most 148 of 164, bias at most 121 of 136, distinct sum at least 247,335 (bound 245,760) |
| Accepted margins, memory-hard | saturated at most 157, bias at most 126, distinct at least 247,571 |
The 20,000-seed reading holds at 10^5: the stand-ins disagree on a one-bit property (a final register bit pinned in
2,047 of 2,048 hashes against 2,048) and on four programs' output bias within sampling noise, never on saturation
or the distinct count, and never in the attacker's favour. Gate: the 39 edge disagreements reproduced and bounded.
Sub-row (a): PASS.
### (b) The hot-set search, 10^6 seeds
| Quantity | 1,000,000 seeds |
|---|---|
| Programs with any hot bucket (strict) | 75,400 (7.5 percent) |
| Programs flagged at the gate reading (hot share at least 1 percent of reads, or 7 constant address bits) | 11,696 (1.17 percent) |
| Worst hot share | 17.3 percent of the program's reads (seed 842871) |
| Hot-share bands (block init) | 0: 932,106; under 0.1 percent: 11,721; under 0.5: 3,645; under 1: 41,581; 1 percent and over: 10,947 |
| Sites with 7 or more constant address bits | 0 (max 6) |
| Fewest distinct addresses at a site in 2,048 evaluations | 43 |
| Most evaluations reading one address at a site | 1,838 of 2,048 |
| Mean hot share over programs | 0.063 percent |
Gate: zero passing programs with a hot set under 1 percent of items. FAIL on the class v4 stream by the letter:
11,696 passing programs concentrate 1 percent or more of their reads on a hot set. The mechanism is F9-1 above,
the or-saturated load source, which is the same fault class the F8 row found from the cross-hash histogram
(AP-F8-1: a load whose source's last writer is lossy); the two harnesses found it independently, one from the
program's address trace per site, one from the item histogram across hashes. F9-1 therefore merges into AP-F8-1,
and the fix is the amendment shipping in 0.3.20 (a load's source drawn only from registers whose last writer injects
or is a rotate). Sub-row (b): FINDING (AP-F8-1 class); re-gated on the amended stream with this harness below.
### (d) Exhaustion count on the frozen class v5 tip, 10^5 chain-shaped seeds (1c420786; the 0.3.24 gate line)
Run: igneum-pow class-v5 1c420786 (pairing e5a4ac5978462156, the harness draws through `Epoch::chain_program` with the
era-composed class, `--chain`, the dn3 state leaves), build-1, ten chunks of 10,000 seeds under `lease pool 4` (chunks 1,
3 and 5 to 9 under class v5, chunks 0, 2 and 4 re-leased under class release on the coordinator's order), the binary
copied into `frozen-1c420786-f9/bin`, 14,477 to 14,480 s per chunk (about 1.45 s per seed), last chunk written
02:34:54 UTC on 8 October 2026 (`par0.tsv` to `par9.tsv`, `par0.log`, `par2.log`, `par4.log`). The interim line at
00:55 UTC (seeds drawn to that minute, 0 exhausted, 0 panics, max 30) cleared the 0.3.24 move; this is the record line.
| Quantity | 100,000 seeds |
|---|---|
| Seeds drawn | 100,000 of 100,000 |
| Exhausted (the 256-attempt cap reached, or the deterministic last resort used) | 0 |
| Panics on the draw path (caught and counted) | 0 |
| Past attempt index 31 | 0 |
| Max attempt index | 30 (one seed) |
| First-draw acceptance (index 0) | 31,454 (0.3145) |
| Mean attempt index (draws per seed) | 2.185 (3.185) |
| Index 8 or above | 4,862 (4.86 percent) |
| Index 16 or above | 255 (0.255 percent) |
Histogram by attempt index: 0: 31,454; 1: 21,460; 2: 14,660; 3: 10,263; 4: 7,047; 5: 4,701; 6: 3,297; 7: 2,256;
8: 1,532; 9: 1,027; 10: 702; 11: 509; 12: 365; 13: 216; 14: 153; 15: 103; 16: 80; 17: 56; 18: 39; 19: 20; 20: 24;
21: 9; 22: 6; 23: 9; 24: 3; 25: 4; 26: 2; 27: 1; 29: 1; 30: 1. The tail is geometric at a ratio of about 0.69 per
index (the per-draw rejection rate under (a'), (c'), (c'') and (c''') together), so the chance of reaching the cap at
256 is far below the count; the observed bound is under 3e-5 per seed at 95 percent (0 of 10^5).
Verdict: PASS. Consequences per tier: no epoch seed in 10^5 fails to draw, so the AP-F8-2 liveness halt (a stuck
epoch for a node operator, a dead epoch for a miner, a paused chain for a holder) has no observed case on the frozen
tip; the draw costs every node about 3.2 candidates per epoch, which at about 1.45 s per seed on one box core is under
five seconds of CPU per epoch and does not change any tier's cost.

View file

@ -0,0 +1,53 @@
# O-1.14 bench start 2026-10-07T08:49:03Z host root@ssh9.vast.ai
Warning: Permanently added '[ssh9.vast.ai]:35608' (ED25519) to the list of known hosts.
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
# binary copied, sha256 6d28678359d3d6bd158b245f7e522d6f2a5b0704d9997c0fb50ebc3471a9ebe5
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
# cpu: Intel(R) Core(TM) i7-9700K CPU @ 3.60GHz
# cores: 8 mem: 31 GB
# glibc: ldd (Ubuntu GLIBC 2.39-0ubuntu8.9) 2.39
# clock MHz: 4169.856
08:49:07 up 20 days, 10:27, 0 user, load average: 0.13, 0.07, 0.05
## --class v2 08:49:07Z
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
cache: fill 283.0 ms on one core (2^26 words, 256 MiB, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 48c4f5bf24166b2e
warp base 0: single cold run 1.582 ms, 4096 items derived, lane0 42246ba99fc58e4f lane31 b08446b1f2de7793
warp base 4096: single cold run 1.392 ms, 4096 items derived, lane0 3d3903e310ca038f lane31 61c242509efdccdd
warp base 1000000: single cold run 1.374 ms, 4096 items derived, lane0 f218c1bd58e6dfe0 lane31 6c3b2c11adfbfcac
CPU verify: 1.280 ms per 32-lane warp, avg of 50 (checksum 19297e99c7b9a55e)
## --class mx8 08:49:09Z
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
cache: fill 276.6 ms on one core (2^26 words, 256 MiB, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 48c4f5bf24166b2e
warp base 0: single cold run 5.394 ms, 4096 items derived, lane0 19b56348bc85304d lane31 359192708e4f754a
warp base 4096: single cold run 5.285 ms, 4096 items derived, lane0 62fb132a9943127a lane31 7d7866cb9cfca8ff
warp base 1000000: single cold run 5.293 ms, 4096 items derived, lane0 86b6cb0e13d89b03 lane31 9c004678515e44ec
CPU verify: 5.267 ms per 32-lane warp, avg of 50 (checksum 653a23f7ee1c8c63)
## --program-class v4 08:49:11Z
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
cache: fill 276.6 ms on one core (2^26 words, 256 MiB, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 48c4f5bf24166b2e
warp base 0: single cold run 6.334 ms, 4096 items derived, lane0 2576769ee4a14c8d lane31 c58ddcb717dd3370
warp base 4096: single cold run 6.198 ms, 4096 items derived, lane0 1ce77a600ec573b4 lane31 03600a05ffba0055
warp base 1000000: single cold run 6.174 ms, 4096 items derived, lane0 6b390e64bbdd91ce lane31 91c944d603539c62
CPU verify: 6.006 ms per 32-lane warp, avg of 50 (checksum 17e36e7905b81375)
## --class dr368 08:49:14Z
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
cache: fill 276.3 ms on one core (2^26 words, 256 MiB, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 48c4f5bf24166b2e
warp base 0: single cold run 5.540 ms, 4096 items derived, lane0 c77c7625bbe0f452 lane31 c8e84ff2655d9934
warp base 4096: single cold run 5.433 ms, 4096 items derived, lane0 7229bd981a5786ca lane31 1c5495083ae60453
warp base 1000000: single cold run 5.430 ms, 4096 items derived, lane0 5533769c9cdbf0a7 lane31 2426905704457b11
CPU verify: 5.426 ms per 32-lane warp, avg of 50 (checksum 94fcbf0a77e03bdc)
## --class dr736 08:49:16Z
Welcome to vast.ai. If authentication fails, try again after a few seconds, and double check your ssh key.
Have fun!
cache: fill 275.7 ms on one core (2^26 words, 256 MiB, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 48c4f5bf24166b2e
warp base 0: single cold run 10.290 ms, 4096 items derived, lane0 e23d389f3eea0c83 lane31 6605db059b381bd9
warp base 4096: single cold run 10.072 ms, 4096 items derived, lane0 fdb4b214da8ce292 lane31 db698d03437d74f7
warp base 1000000: single cold run 10.056 ms, 4096 items derived, lane0 534671b1bf5cea36 lane31 733b123353f13d21
CPU verify: 10.042 ms per 32-lane warp, avg of 50 (checksum bf79909c25836153)
# O-1.14 bench end 2026-10-07T08:49:19Z

View file

@ -1,6 +1,6 @@
# Block rate on Devnet 2: 10 blocks per second against 1, on the rented fleet
6 October 2026, the project lead's experiment (through the coordinator, 17:5xZ): "Devnet 2 at a higher block rate for solo miners
6 October 2026, the founder's experiment (through the coordinator, 17:5xZ): "Devnet 2 at a higher block rate for solo miners
(Kaspa's answer)". Branch `gpu-fleet`; the node profile on the fork branch `devnet2-bps` (1279a1d6 on release-0.3.14-node
4c6b129d): a suffixed devnet started with `IGNEUMD_DEVNET_BPS=10` (or 5) runs `BlockrateParams::new::<10>()` with the
TenBps subsidy and target time, Kaspa's Crescendo-style constants for that rate (k 124, merge depth, sample rates,

View file

@ -42,14 +42,14 @@ F8's proposed gate (the top 0.1 percent within 1.2x of the window-model control
A hot-set cache of the top 0.1 percent of items is 16,777 items x 64 B = 1.07 MB of SRAM, 0.53 mm^2 and $0.25 at 0.49 mm^2 and $0.23 per MB. It serves 0.52 percent of p1's reads (0.16 of them the window model's), 4.6 percent of p3's. The hash is latency-bound on its dependent reads, so a read served on die is time saved: a chip gains at most 1.005x on p1 and 1.048x on p3 from the cache. The ceiling under the live rule: part (c)'s 120-of-128 floor admits one site repeating its item in all 8 iterations and no more (two saturated sites fail it), so at most 8 of 128 reads, 6.25 percent, can sit on a constant item, and a chip's edge from this whole class is at most 1 / (1 - 0.0625) = 1.067x, in 64 bytes of SRAM, on the hours whose program carries such a site. The public claim rests on 2x margins (chip-model-v3.md); 1.067x does not move it, and the union of the windows is still the whole dataset every hour, so no window-level cache exists. What moves: per tier nothing in rate or watts (the honest card reads the hot item from L2 as the chip would), and the 5 percent rule of 2.0 is untouched.
## 5. The two options for the flip, priced (main's ruling 3; nothing ships on this without the project lead's word)
## 5. The two options for the flip, priced (main's ruling 3; nothing ships on this without the founder's word)
| Option | What changes | Cost | Risk |
|---|---|---|---|
| A. A class amendment in 0.3.19 before the flip: the generator draws a load's source from the registers whose last writer injects (or rule (a) tightened to the same), class v4 re-pinned | a new program stream: new vectors, the seven gate packs re-exported, the six gates again (the hash side G1 to G3 and the verifier re-run here in about an hour of Mac and PC 2 time; G4 to G6 the node lane), every node before the flip by the one-box-at-a-time fleet rule | hours of gate time, a fleet rollout, the 0.3.19 ship on the line | a node that misses the build splits the chain at the flip; the fix itself is small (one draw rule) |
| B. Hold v4 at the floor as it is; the source rule in class v5 | nothing on the devnet; the attack-pass record carries the window null and the bound | a hot set on 48 percent of hours worth up to 1.005x to a chip, on 5 percent of hours up to 1.05x, 1.067x at the rule's ceiling, no chain risk | the public line must state the bound, not "uniform" |
The number that decides it: 1.067x at the ceiling against the 2x margin of the chip claim. Recommendation: B, with the v5 item below, unless the project lead wants the tail tight now.
The number that decides it: 1.067x at the ceiling against the 2x margin of the chip claim. Recommendation: B, with the v5 item below, unless the founder wants the tail tight now.
## 6. The acceptance bound for the next class (main's ruling 4)

View file

@ -69,7 +69,7 @@ Proving is a separate budget (the 15.6 GB peak the 12 GB mine-and-prove question
## 4. Table 3: the public sentences against the numbers
| Where | Sentence now | What the tables give | Proposed sentence (the project lead decides the wording) |
| Where | Sentence now | What the tables give | Proposed sentence (the founder decides the wording) |
|---|---|---|---|
| `site/index.html` 443 | Memory: "2 GB, fixed" (RandomX) / "2 GB, growing" (Igneum) | 2 GiB at genesis, plus 0.5 GiB a year on average under either option | "2 GB, growing 0.5 GB a year". The row is right; the rate is the useful addition |
| `site/index.html` 461 | "Any 4 GB card, approximate." | True at genesis (2,584 to 2,834 MiB of a 3,072 MiB budget). Ends at 1 to 1.5 years under (a), year 4 under (b) | "Any 4 GB card at launch, 8 GB for the long run, approximate." |

View file

@ -202,7 +202,11 @@ the `f = 1` chip: it is the 55 W without the 271.
### 5.4 The curve
Per row: reads per hash = 128 f; items recomputed = 128 (1 - f); ops per hash = 128 (1 - f) x 9,360 + 512.
Correction, 7 October 2026 (the in-house adversarial pass, lane adv-cache-3, report 091edc34): the partial-store rows above and adv-cache's Q1b table price a chip that holds every k-th line of the 64-line chain and recomputes a read at offset o in o evaluations ((k - 1) / 2 on average). The exact pebbling optimum for the chain (dynamic programming, checked against exhaustive search at 10 to 16 lines) sits under that curve: blocks per read 16.0 against 31.5 at f = 1/64 (the one held line belongs at line 32, not line 0), 10.5 against 15.5 at 2/64, 6.09 against 7.5 at 4/64, 3.17 against 3.5 at 8/64, 1.45 against 1.5 at 16/64, equal from f = 1/2. So a chip holding 1/64 of the cache pays 9.3x the item's ops, not 17.4x; at f = 1/2 and above nothing moves, and the SRAM column and the full-store verdict stand (no point on the curve beats the full store under the op budget or under energy).
Memory-bound rate = the ceiling / (128 f). Compute-bound rate = 50 T op/s / ops per hash (the section 1 budget).
Second correction, 8 October 2026 (the in-house adversarial pass, lane adv-cache-2, report section 2.3 and its Q3(3) window-layer reading, tip 3f50d6c4): the partial-store rows model a chip that holds a fraction f of the ITEMS chosen uniformly, so it serves f of the reads and recomputes 1 - f. The item-read distribution is not uniform: the exact window-layer distribution (matching the 4,096-program census to four digits; top quarter mean 0.3382, top half 0.5811 of reads) lets a chip holding the hottest f of items serve 0.4219, 0.7188 and 0.8907 of reads at f = 0.25, 0.5 and 0.75 on the measured programs, so the f = 0.25 and f = 0.5 rows overstate the recompute share by up to 2.3x (1.8x on the first shard's read) and the f = 0.75 row by about 2.3x on the miss side. The f = 1 row, the SRAM column and the full-store verdict do not move (a chip that holds everything recomputes nothing either way), and no served number rests on f under 1; the partial-store rows stay as the uniform-store bound with this note until the hottest-f rows are drawn from the window distribution, which is the next pass of this section.
The rate is the smaller; "binding" names it. Power = rate x (128 f x E_read + 128 (1 - f) x 6.3 nJ) + static (memory,
controller, and 20 W for the recompute die's clocks and leakage when `f < 1`). Energy per hash = power / rate. "Gain,
rate" = rate / 136.1 MH/s (per chip, the section 2 metric); "with the 3x factor" multiplies the compute-bound rate by
@ -294,7 +298,7 @@ low on this result; it stays a reserve family. What does move the `f = 1` rows,
| Read granularity | The chip pays the same 32-byte atom the 5090 pays; the 9070 XT pays 64. Wider honest reads (w16, measured, layer 1) give the chip nothing and the 5090 nothing; w64 made the 5090 bandwidth-bound (71.9 MH/s) | w64 costs the 5090 47 percent | Not a lever; the decision to stay at 4 B stands |
| Latency | A longer chain (more reads per hash) scales the chip's rate and the card's rate together; lane state is 64 B, so lanes are free to the chip | Nothing per se | Not a lever: the rate per chip is lanes / latency on both sides and the chip has more lanes per watt |
| The denominator: the 5090's watts at the hash | The gain is 2.40 microjoules over the chip's 0.47; the card's 326 W is 92.9 percent utilisation spinning on loads. At a 250 W cap holding 136.1 MH/s the gain reads 3.9x (GDDR7) and 5.7x (HBM3); at 200 W, 3.2x and 4.6x | None if the rate holds under the cap; the measurement is one PC 2 job (`nvidia-smi -pl 200, 250, 326`, two minutes each, STATUS lines as the rate) | The first measurement to run; it moves every row and costs nothing. Owed (PC 1 is not released; PC 2's budget is the coordinator's) |
| Program work in the latency shadow | The hash hides 512 ops per hash behind 128 reads; the 5090 could hide 330,000 (45.2 T / 136.1 M) before compute binds, the M5 Max about 290,000 and the 9070 XT about 650,000 (their ALU budgets approximate, from memory). Work in the shadow is free in hash rate and costs the card watts it now wastes: at N ops per hash the card rises from 326 toward 575 W (linear, approximate) and the chip must add a core that runs the per-epoch random program, at k times the GPU's 5.5 pJ per op (the 5090's marginal ALU energy, (575 - 326) / 45.2 T). At N = 100,000: the card 401 W, 2.95 microjoules; the chip 1.02 at k = 1, 0.83 at k = 1.5; gain 2.9x and 3.5x. At N = 200,000: 477 W, 3.50; chip 1.57 and 1.20; gain 2.2x and 2.9x. At N = 330,000: 575 W, 4.22; chip 2.28 and 1.68; gain 1.85x and 2.5x | Hash rate none while every card stays latency-bound (under about 290,000 on the M5 Max); watts up to TGP; the verifier N x 32 ops per warp: 3.2 M at N = 100,000, under 1 ms at the 18 G op/s the x8 verifier shows (38 M ops in 2.08 ms), inside the 10 ms gate; `INSTR_COUNT` and `ITERATIONS` are prototype values to be fixed at gate 1 (spec 1.4) | The only lever that moves the f = 1 row toward 2x, and only if the chip's core is no better than a GPU's on a random program (k near 1: RandomX's argument, and the project lead's goal in the brief's words, "build a better GPU than NVIDIA"). It reaches 1.85x at the 5090's full ALU budget and k = 1, not under; combined with a 250 W cap it reads about 1.4x (approximate). It is item 2's idea applied to the program, not to the item derivation |
| Program work in the latency shadow | The hash hides 512 ops per hash behind 128 reads; the 5090 could hide 330,000 (45.2 T / 136.1 M) before compute binds, the M5 Max about 290,000 and the 9070 XT about 650,000 (their ALU budgets approximate, from memory). Work in the shadow is free in hash rate and costs the card watts it now wastes: at N ops per hash the card rises from 326 toward 575 W (linear, approximate) and the chip must add a core that runs the per-epoch random program, at k times the GPU's 5.5 pJ per op (the 5090's marginal ALU energy, (575 - 326) / 45.2 T). At N = 100,000: the card 401 W, 2.95 microjoules; the chip 1.02 at k = 1, 0.83 at k = 1.5; gain 2.9x and 3.5x. At N = 200,000: 477 W, 3.50; chip 1.57 and 1.20; gain 2.2x and 2.9x. At N = 330,000: 575 W, 4.22; chip 2.28 and 1.68; gain 1.85x and 2.5x | Hash rate none while every card stays latency-bound (under about 290,000 on the M5 Max); watts up to TGP; the verifier N x 32 ops per warp: 3.2 M at N = 100,000, under 1 ms at the 18 G op/s the x8 verifier shows (38 M ops in 2.08 ms), inside the 10 ms gate; `INSTR_COUNT` and `ITERATIONS` are prototype values to be fixed at gate 1 (spec 1.4) | The only lever that moves the f = 1 row toward 2x, and only if the chip's core is no better than a GPU's on a random program (k near 1: RandomX's argument, and the founder's goal in the brief's words, "build a better GPU than NVIDIA"). It reaches 1.85x at the 5090's full ALU budget and k = 1, not under; combined with a 250 W cap it reads about 1.4x (approximate). It is item 2's idea applied to the program, not to the item derivation |
| The clock (item 4) | The f = 1 chip is a commodity-memory controller project: by the Ethash precedent, 32 months to a first chip at the largest prize, and a chip over 2x at 65 months | None | The issuance trigger and the share-pattern detector matter more than any item-derivation change |
So: item 2 can wait; the power-cap measurement runs first; the program-length lever is the Counter ASIC 3.0 design
@ -335,6 +339,34 @@ measurements land.
- The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000
program ops are estimates; the program-length lever is a design item with its own measurements, not a result.
### 5.10 Class v5 on, the shadow at zero: does the state-derived dataset make the shadow unnecessary? (7 October 2026, 21:3x UK, the founder's question "we need a solution, deep research, other methods")
The question: with class v5 (the dataset built from the chain's execution state, refreshed per window) and the latency-shadow work of class v4 set to zero (class v3 energy on every GPU), what edge does the strongest chip keep over an RTX 5090 per joule? If it were at or under about 2x the shadow could come off after v5 and the class v4 premium (145 W on a 5090 at the unlocked core, 88 W at the 1,400 MHz lock, measured 7 October 2026) would vanish.
What class v5 changes for the chip, from `docs/design/class-v5-stored-state.md` sections 2 and 2a: the recompute chip (`f = 0`, the dataset derived on the fly from the cache) and the stateless or stale chip are removed as categories, because every item takes a leaf of the state and the leaves refresh every window. What it does NOT change: the strongest chip was never one of those. It is the `f = 1` stored-dataset chip of 5.5, a GPU's memory system without the GPU, and under class v5 it needs one thing more, the window's leaves, which section 2a.2 prices honestly: one node serves a whole farm, the leaves ship at 16.5 KB/s to 10,000 members today (45 MB per member per window over a 1 Gbit/s WAN before the state is 7,500x today's), and the rebuild on the chip is the same 32 ms per window every GPU pays. The dataset is still derived from a seed and the state, so a central node compresses everything but the state's bytes, and the chip stores the result as before.
The arithmetic, on the model's own figures (5.3: the activate-bound ceilings, 2.0 nJ per random read on GDDR7 and 1.2 nJ on HBM3, the static and controller watts; the 5090 at 136.1 MH/s on 326 W, 0.417 MH/W; 128 loads per hash, the shadow at zero so no ALU beside the memory). The node is a desktop-class CPU with an NVMe and 32 GB at about 85 W (approximate, from memory of such machines; a full node with the EVM executor at 1 block/s), shared by a farm (100 chips: 0.85 W each) or carried by every chip (the attacker's worst case, 85 W each).
| Chip, class v5 on, shadow at zero | MH/s (model) | W with a farm-shared node (0.85 W) | MH/W | Edge over the 5090 per joule | W with a node per chip (85 W) | Edge |
|---|---|---|---|---|---|---|
| GDDR7 `f = 1`, 16 devices (the 5090's own memory without the GPU) | 166 | 78 | 2.12 | 5.1x | 163 | 2.5x |
| HBM3 `f = 1`, one stack | 84 | 28 | 3.02 | 7.2x | 112 | 1.8x |
| HBM3 `f = 1`, eight stacks (an H100-class package) | 666 | 175 | 3.80 | 9.1x | 259 | 6.2x |
Every figure is modelled (arithmetic on cited memory figures, approximate where 5.3 marks it); none is measured; the node's watts are an approximate from memory.
Reading: NO. Class v5 with the shadow at zero leaves the strongest chip at 5.1x (GDDR7) to 9.1x (HBM3, eight stacks) per joule, the class v3 figures of 5.6 less a rounding, because the node is a farm cost and not a chip cost; only a chip forced to carry its own node falls near 2x, and only the small ones (one HBM3 stack at 1.8x, the GDDR7 board at 2.5x), while the eight-stack package stays at 6.2x even with a node per chip. So the shadow (class v4's 100,000 ops per hash in the memory wait, which brings the chip to 2.1x at k = 1 and 3.9x on the claimed X9 core) stays the only lever in this model that reaches the memory-system chip, and the class v4 premium is the price of that lever on today's GPUs. What class v5 buys is different and real: the recompute chip and the stale chip are gone as categories, every miner must hold and follow the chain, and a chip's dataset is wrong the moment its node is. The premium itself has two measured levers tonight: the core-clock lock (57 of 145 W back on the 5090 at 1,400 MHz for 1.4 percent of rate; the knee below 1,400 is the second pass's) and the per-card tune the app lands by itself; what would remove it is a shadow whose work is cheaper per op on a GPU than on a chip core (the research lane's question: a shadow shaped for the GPU's idle datapath at low clock, or a memory-side cost the chip cannot amortise), not the state-derived dataset.
Per tier: a miner on class v4 pays the premium and gets the 2.1x to 3.9x chip ceiling in exchange; on class v5 with the shadow kept the ceiling stays and the dataset is the chain's; on class v5 with the shadow dropped the premium goes and the ceiling returns to 5x to 9x. The decision is the founder's; this section gives the number.
### 5.11 Two columns from the counter-asic-4 research file (7 October 2026, 22:3x UK; `docs/analysis/counter-asic-4-research.md` sections 15 to 19 at bca23f96, the research lane's reading, carried here as the chip model's own columns)
The k column. k is the chip core's energy per op over the GPU's at the same operating point, and the model's 3.4x row takes k about 0.33 for an ALU-shaped core. On the public figures the int8 tensor tile is the one GPU block whose energy per op a chip at the same node cannot undercut with certainty: the 4090 measures 0.056 pJ per MAC; NVIDIA's 5 nm INT4 test chip reads 0.021 pJ per MAC at 0.46 V and about 0.1 at nominal (JSSC 2023, via Dally's NASEM slides; claimed), so INT8 at 2x to 4x that gives k 0.7 to 3 with the centre near 1; every ALU-shaped block reads k 0.3 to 0.8 on the same sources. A shadow built of tensor tiles at the ALU shadow's premium (about 11,400 u8 tiles per hash) therefore gives 2.1x at k = 1 and 1.6x at k = 1.5 and removes the k 0.3 column from the table; it needs a SIMD byte-dot verifier (the scalar one at 12.4 ms fails the 10 ms gate). This is a design candidate, not the shipped stream: the shipped shadow is ALU-shaped and its row stays 2.1x at k = 1 and 3.4x at k about 0.33.
Correction, 8 October 2026 (the research lane's microbench on the 5090, counter-asic-4-research.md 15.1a and the corrected 20.3 and 20.4 at 71fd465b): the per-MAC figures above are wrong by a factor of 32. A `mma.m8n8k16` tile is 1,024 multiply-adds per warp, 32 per lane, so a hash does 32 MACs per tile, not 1,024; the 4090's "0.056 pJ per MAC" is 1.8 pJ, and the 5090 at the ALU shadow's premium reads 2.9 pJ per MAC unlocked and 1.5 pJ at the 1,300 MHz lock (the packs job, 366,080 MACs per hash; the microbench's dependent u8 tile 4.1 and 2.2, the wide s8 m16n8k32 tile 1.36 and 0.83). Against the same 5 nm MAC array figures (0.04 to 0.4 pJ per INT8-class MAC, claimed) a chip's k on tile work is therefore 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, not near 1: at the same premium a tensor-shaped shadow leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. The tensor-tile column (2.1x at k = 1, 1.6x at k = 1.5) is withdrawn as a candidate; its premise, that a chip's MAC is no cheaper than the GPU's, is false by 4x to 30x on the public figures. The shipped row is unchanged: 2.1x at k = 1 and 3.4x at k about 0.33, the ALU shadow at the operating point's knee, measured four times at 82 to 90 W.
The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure, the 16 x 27 placement, was closed on 7 October 2026 at night when it failed the value-level acceptance test across drawn eras, so the 200 M row rests on no construction shown to exist until a sound per-load class, one pass of a 432-instruction sub-block per load, is drawn, accepted and measured). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job.
## 6. The per-day derivation (item 2)
6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything

View file

@ -110,7 +110,7 @@ run drops from about three minutes to about one. The hosted runner then serves o
pipeline (MSVC, WebView2, Inno Setup, PowerShell 5.1, which a Linux box cannot provide). The flip is main's call after the
0.3.15 cut, per build-server.md section 7.1.
## 6. The build box's own red rows (the project lead, 22:3x UK: "also make sure we are fixing and learning from all the errors here")
## 6. The build box's own red rows (the founder, 22:3x UK: "also make sure we are fixing and learning from all the errors here")
Source: `/srv/builds/_log/builds.jsonl` on igneum-build-1, 139 rows from the first build at 18:38 UK to 22:11 UK on
6 October; 34 with a non-zero exit. Until this change the box kept no output of a run (it streamed to the agent's terminal),

View file

@ -0,0 +1,232 @@
# Class v6, the family gate: how a parameter family is cryptanalysed as a family
Research lane D of the class v6 rotating-family design, under the Counter ASIC coordinator. First cut 8 October 2026, 11:0x to 17:00 BST (the deliverable clock); the full report by 09:00 BST on 9 October. Branch `class-v6-family-gate` (this document) and the harness branch `family-gate-v5` (the class v5 crate at 8f481459 plus the census harness, never a chain path). Internal research, not an independent review; nothing here touches a served number, the devnet or the testnet object; no consensus code this week. Every figure carries its source: a lane report on the mirror, a log on a box, or an arithmetic step shown in the text. Figures from memory are labelled approximate.
The question. Class v6 (the founder's word at 11:1x BST, 8 October 2026) draws per era the parameters a release used to fix (layer 1), lets the dataset track the chain state (layer 2), rotates instruction families by height (layer 3) and generalises the acceptance floor and the uniformity test to each era's draw with a redraw on failure (layer 4). A chip is then built against a FAMILY of hashes, not one hash. The in-house pass of 7 October and the attack board F1 to F10 bounded ONE member of the family (class v4 sub-version 3 at 017e7037 and class v5 at 1c420786). This document says how the family is bounded: what the space is, how many eras must be drawn to say anything about the worst era, which tests the chain runs itself on every era and epoch, what the offline board looks like over the testnet period, what a published proof of testing contains, and what the whole apparatus cannot see.
## 0. One page
| Item | The line |
|---|---|
| The space (section 1) | Seven drawn axes with a tested band each: the mixer multiplier `m` in {4, 8, 16} (8 and 4 measured on cards and the verifier; 16 the verifier's open row), the ten non-load op weights within B points of table 1.4.2 with shuffle and mulhi never raised (B = 2 proposed in spec 1.13.1, B = 4 the research lane's proposal), the read width in {1, 4} words (w16 measured 5 October, w64 excluded), the shadow block shape in {64, 128, 256} instructions at 6,912 per iteration (measured 6 October), the shadow placement fixed at the whole block after instruction 63 (the per-load placement is a known-failed corner: 1.4 percent acceptance, 42 of 64 seeds exhausting 32 attempts), the program length N by the ladder's signal inside rungs 0 to 2 (rung 3 inadmissible by F6), and the existing era draws (the stride `M`, the rotation `R` in 1..31, the interleave, the per-site windows). The day key (ROT, MUL, RC) is drawn daily under the AP-F4-1 redraw rule |
| Which tests depend on which axis (section 1.3) | The acceptance rule and every index statistic ((c''), (c'''), the bucket bound, the bit-bias read, the attempts census) run on the closed-form dataset and never see `m`: the mixer axis is bounded per family by adv-mixer-3's rows and F4, not per era. The width axis changes the index space of every site ((c'')'s expectation `E_s` must divide the window by the width, a spec change the harness carries). The shape and weight axes change the draw's stream and so every per-era statistic |
| The sampling bound (section 2) | With `n` eras drawn from the chain's own distribution and every one passing a per-era test, the fraction of eras that would fail it is under `3/n` at 95 percent confidence (the rule of three; Wilks 1941 for the order-statistic form): 190 eras bound the failing fraction at 1 in 64 (one era in 32 years of 180-day eras), 3,067 at 2^-10. The bound is on the TESTED property only; the union bound over the per-era tests' own false-pass rates gives the family's false-pass rate, and each test's false-pass rate is itself a rule-of-three number on its known-failed cases (the (c''') floor refused 9 of 9 live hot sets, so its miss rate on that class is under 0.33 at 95 percent: nine cases bound nothing tighter) |
| The per-era tests (section 3) | Three rings. Ring A, the chain runs per era at the cut, microseconds: the structural redraw rules (the band membership of every draw, the day-key cost rule, the product-bias rotation class recorded). Ring B, the chain runs per epoch, seconds: the acceptance rule generalised to the family's shape ((a) (b) (a') (c) (c') (c'') (c''') plus the two next-class tests, the per-site largest-256-item-bucket bound and the index-bit bias read, on the same 2^20 pass). Ring C, the gate runs offline per drawn era, box-hours: the F8-form census at 2^24 on 64 seeds, the attempts census, the exhaustion count. Each row carries its known-failed shape and its clean-seed cost so the redraw is priced |
| The board (section 4) | F1 to F10 and the nine lanes classified per-era, per-family or chain-run, with the harness taking the era draw as an input; the proof of testing is one JSON record per drawn era (the draw, the object, the binary, the seeds, every verdict with its statistic, threshold and log sha256, the core-seconds) and one family summary (eras per stratum, the worst era per test, the refuse rate per band, the bound's arithmetic) |
| The honest line (section 5) | The gate sees what its tests see. It does not see the diffuse era-stride class below the bit-bias read's resolution (a bias under 0.3 percent at 2^20), a mechanism that lives in a joint corner no stratum names, a mixer weakness below adv-mixer-3's bands, or an era whose draw sits in the untested 1 in 64. A GPU-like chip loses nothing on any era; what a corner buys it is a per-era hot set bounded at 1.002x (the floor) or 1.0024x (the diffuse class), or, at `m = 4`, a mixer with 2 applications of measured margin instead of 6 |
| Running now (section 6) | The harness `family_gate_era_census` on build-1 from 11:16 BST: one drawn era per seed, every family parameter from the era's own stream, the chain draw through the real rule keyed on the family's shape, one TSV row per era with the (c'') ratio, the bucket ratio, the bit bias and the attempts; the base control (`v5_attempts_census`, the shipped draw) beside it. The coverage table in section 6 is filled from the logs at 16:xx BST |
## 1. The parameter space
### 1.1 The axes, their bands and the margin each band was read from
| Axis | Band proposed for the draw | What fixed the band (the measured rows) | The tested margin inside the band | Known-failed corner, named |
|---|---|---|---|---|
| Mixer multiplier `m` (applications per round, `LoadClass::mixer_mult`; 72 per item at `m = 8`, 8 between dependent reads) | {4, 8, 16}; the research lane's outline; this lane's reading: {8} plus {4} as a corner, {16} only after the verifier row | x4 and x8 measured (`docs/plans/mixer-x4.md` section 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the daily build 42 ms at x8 on the 5090, measured); x16 estimated at about 3.7 ms per unit on the M5 Max core and 9 ms on a 2019-class core (chip-model-v3 section 3 item 1), against the 10 ms verifier gate | adv-mixer-3 (`report-mixer-3.md`, "The round margin, stated"): no statistic survives 2 of the 8 keyed applications between reads at bands 0.0015 (Q2, 2^24 states), 0.00026 (2^27), 0.00018 (Q2b, 2^28); SAT solved at k = 1 in 137 s and timed out at k = 2, 3, 4; two real days and eight random days agree. So at `m = 8` 6 of 8 applications are margin; at `m = 4` the same rows leave 2 of 4; at `m = 16` 14 of 16 | None measured. The corner is `m = 4` by margin and `m = 16` by the verifier; the all-ROT-equal day key is the day-level corner (rejected by the AP-F4-1 rule) |
| Op-mix weights (the ten non-load families, sum 75) | Each weight within B points of table 1.4.2 (add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4), renormalised by largest remainder; shfl and mulhi never raised; B = 2 (spec 1.13.1, proposed) or B = 4 (the research lane) | The microbench (`counter-asic-4-research.md` 15.1a, the 5090 alone, 8 October): pJ per counted op at the stock clock and the 1,300 MHz lock: arx 11.3 / 6.2, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4; the shuffle is 4.9x the add, so a draw that raises shfl raises the card's premium per instruction with no better `k` for a chip (15.1b). F7's era census: the op-weight corners (15 to 31 of 75) read 1.0x against the GPU, 0 memory effect | The acceptance's rejection rate moves with the lossy share: at the table 0.68 per candidate ((a') 83.5 percent of rejections; adv-accept row 90, 20,000 seeds); the F7 census at B = 2 read every corner clean for the chip; nothing was measured at B = 4 before today (section 6) | The lossy corner: or, mul and mulhi all at +B. It raises (a') rejections and so the attempts per seed; the exhaustion bound `r^256` is the number to read there (section 3, ring C) |
| Read width `W` (words per load) | {1, 4} | `docs/plans/read-width.md` (5 October): w16 within the rules (5090 139.8 against 136.1 MH/s, 9070 XT 17.90 against 18.15, M5 Max within 1 percent); w64 bandwidth-bound (5090 share 0.58, 71.9 MH/s): out; the per-load mixes out on the 5 percent spread rule | The acceptance runs unchanged on aligned addresses (read-width.md section 2), but (c'') and (c''') read a site's distinct ALIGNED indices against `E_s = N - N^2 / 2W_s`: at `W = 4` the index space is `W_s / 4`, so the expectation must follow or every width-4 program reads 0.976 at a quarter window and is refused by the 0.98 floor. The harness divides the window by the width (`site_window_words`); the spec line is owed | `W = 16` (w64): excluded by measurement, never in the band |
| Shadow block shape (instructions per block, passes per iteration) | {64 x 108, 128 x 54, 256 x 27}: 6,912 shadow instructions per iteration at rung 0 | 64-instruction blocks ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (6 October, measured); 1,024 cost the M5 Max 17 percent (out) | F1's shadow redundancy bound (3.0 percent of peephole-removable instructions, the v5 generator's `SHADOW_REMOVABLE_MAX_PERMILLE`) was set from F1's census at 256 instructions; the removable fraction per block is a property of the block's length, so the fraction must be re-read at 64 and 128 (a per-family row). The acceptance's (a') fixpoint and the shadow's execution in (c) take any length | The per-load placement (16 sub-blocks of 16 after each load): dead as a chain class (`counter-asic-4-research.md` 20.2a-close): 22 of 1,621 candidates accepted, 42 of 64 seeds exhaust 32 attempts, candidate 0 carries index bit 0 set in 40 of 1,024 addresses at site 0 (z 29.5). The placement axis therefore has ONE value in the band |
| Program length `N` (counted ops per hash) | Not drawn: the ladder's signal (rungs 0 to 2); the research lane's outline keeps it so | F6: the worst of 10^5 class v4 programs 8.708 ms on the half-core proxy, so N about 300,000 fits under 10 ms (rung 2 admissible, rung 3 not); the M5 Max loses 3.3 to 10 points across the band (measured ladder) | The verifier's worst case was censused at rung 0 only; a draw inside the ladder needs the 10^5-program worst case at the top rung (per-family row) | Rung 3 (inadmissible by F6) |
| Era stride and interleave (`M` odd, `R` in 1..31, `pos` four ascending bit positions, the fold rotations, the per-site windows `k_off`, `o`) | As spec 1.13.1 draws them today | F7's 2^24 era census (stride bijective, R, pos, M uniform, no class over 1.1x); adv-cache-2 section 2.3: the same 32 base programs read 2 of 32 with a site over 1.04x under `R = 29` and 16 of 32 under drawn eras (8 over 1.2x, worst 1.75x), because `rotl(x * M, R)` places a product's biased low bits (P(bit 0) = 1/4) at address bits R and up, inside the 28-bit index unless R is 29 or 30; at D = 29 only R = 30 and 31 cut bit 0 | The whole 64-seed F8 census is a census over drawn eras (each f8 seed carries its own era), so the per-era uniformity verdicts of sub-version 3 and class v5 (60 of 64 and 61 of 64 under 1.2x) already sample this axis; the devnet's clean reading at R = 29 is the accident, the drawn-era figure the chain's | R in 1..27 with a product as the last writer of a load's source (sub-class A); the warp-uniform mulhi source (sub-class B, 1 in 16,000 warps) |
| Dataset size `D` (layer 2) | 2^28 words today, 2^29 at the 2 GiB genesis step, growing on the 1.13.3 schedule with the state as a floor | The acceptance's closed form is pinned at D = 28 (`ACCEPT_DATASET_LOG2`); the live dataset is larger; the stand-in gap measured at 0.0004 under D = 28 (adv-accept Q2, 54 programs) | Unmeasured at D = 29: which R values hide the product bits changes with D, and the window cap `min(k_off, D - 26)` changes the per-site windows | D at a value where the acceptance's D = 28 windows and the live windows disagree: the stand-in gap row of ring C |
| Day key (ROT x 8, MUL x 16, RC x 16, drawn daily from the day seed) | As spec 1.8.4, under the AP-F4-1 redraw rule in class v5 (cost A at most 205 rejected, k >= 1, the eight ROT equal rejected) | F4 and adv-mixer-2: 0 of 2^24 days over 1.1x on both metrics with the rule (8ca66afa); 5.69e-4 of days over 1.1x on LUT area without it, 15 days a century, worst 28 April 2050 at 1.113x | The rule is per day and chain-run; it does not depend on `m` (the cost is per application) | Day 29,337 (redrawn under the rule); the planted mul1all and mulnaf days fire |
### 1.2 The joint space and what a draw covers
Counting the discrete axes alone: 3 (m) x 2 (W) x 3 (shape) x the weight perturbations ((2B + 1)^10 before the caps: about 9.8 million at B = 2, 3.5 billion at B = 4) x 31 (R) x 2^31 (M) x 1,820 (pos) x 31^6 (the fold rotations). The space is of order 2^80 before the per-site windows, which are per program. No census walks it. What a census can do is (i) draw eras from the chain's own distribution (the era stream over random `E_n`) and bound the fraction of eras that fail a test (section 2.1), and (ii) pin the axes that carry a KNOWN mechanism at their corner and draw the rest (section 2.3), so a mechanism that depends on one named axis is read at its worst value. A mechanism that depends on a joint corner nobody named is in the honest line (section 5).
### 1.3 Which tests see which axis
| Test | `m` | weights | `W` | shape | N | era stride, windows | D | day key |
|---|---|---|---|---|---|---|---|---|
| The acceptance rule and the index statistics ((c''), (c'''), bucket, bit bias), on the closed form | no (the mixer never runs) | yes (the stream) | yes (the index space) | yes (the stream, the shadow's execution) | yes (the shadow's passes) | yes (`load_index`) | no (D = 28 pinned) | no |
| The attempts census and the exhaustion bound | no | yes | yes | yes | yes | yes | no | no |
| The F8-form census at 2^24 on the live dataset | yes (the item words) | yes | yes | yes | yes | yes | yes | yes |
| adv-mixer-3's distinguisher rows, F2, F4, adv-mixer-2 | yes | no | no | no | no | no | no | yes |
| F6 verifier worst case | yes | yes | yes | yes | yes | no | yes (the fill) | no |
| F1 shadow redundancy | no | yes | no | yes | yes | no | no | no |
| F3, adv-cache, adv-cache-3 (the chained cache) | no (the cache fill is ChaCha12) | no | no | no | no | no | yes (the cache doublings) | yes |
| F7 era census | no | yes (the perturbation) | yes | no | no | yes | no | no |
| F5 chip model, F10 ladder | the band's top | the band's dearest mix | yes | no | the ladder | no | yes | no |
Reading: the per-era statistical tests are blind to `m` by construction, so the mixer is a per-FAMILY axis with three values and is closed by re-running adv-mixer-3's ladder and F4 at `m = 4` and `m = 16` (section 4), not by drawing eras. The weights, the width and the shape are per-era axes and are what the drawn-era census of section 6 walks.
## 2. The sampling bound
### 2.1 Zero failures in `n` eras (the rule of three, Wilks)
Let a per-era test have a fixed verdict per era and let `p` be the fraction of the family's eras (under the chain's own draw) that fail it. If `n` eras drawn independently from that draw all pass, then with confidence `1 - alpha`, `p <= eps` where `(1 - eps)^n = alpha`, that is `n >= ln(alpha) / ln(1 - eps)`, about `3 / eps` at 95 percent and `4.6 / eps` at 99 percent (Hanley and Lippman-Hand 1983; Wilks 1941 gives the same arithmetic as the one-sided nonparametric tolerance limit on the worst order statistic). The table, with the era length of layer 3 (180 days) turned into years:
| `eps` (the failing fraction bounded) | One failing era expected every | `n` at 95 percent | `n` at 99 percent |
|---|---|---|---|
| 1/16 | 8 years | 45 | 70 |
| 1/64 | 32 years | 190 | 293 |
| 1/256 | 126 years | 766 | 1,177 |
| 2^-10 | 505 years | 3,067 | 4,713 |
| 2^-20 | the F4 and F7 gate's fraction | 3.1 million | 4.8 million |
What this buys and does not. The F4 and F7 gates ask "no class with gain over 1.1x at a fraction over 2^-20" and reach it by censusing 2^24 draws of a CHEAP statistic (microseconds per day key or era seed). A per-era test that costs box-hours (the F8 census at 2^24) can be run on hundreds of eras, so it bounds the failing fraction at the 1/64 to 1/256 class, never 2^-20. That is why layer 4 puts the tests on the chain: a per-era test the chain runs on EVERY era (ring A) or every epoch (ring B) needs no sampling bound at all; the offline census (ring C) is then a check on the chain-run tests' calibration (their clean refuse rate and their known-failed firing per stratum), and a reading of the one statistic the chain cannot afford, the live-dataset census.
### 2.2 The union bound over the per-era tests
Let the chain run tests `T_1 .. T_k` on every era or epoch, and let `delta_i` be the probability that `T_i` passes an era (or program) that carries the failure class `T_i` exists for. The family's false-pass rate on the union of those classes is at most `sum delta_i`. Each `delta_i` is itself bounded by the rule of three on the test's known-failed cases: a test that fired on `f` of `f` known-failed cases has `delta_i <= 3 / f` at 95 percent. The numbers today:
| Test | Known-failed cases it fired on | `delta` at 95 percent | Source |
|---|---|---|---|
| (c''') floor at 0.995 on the few-item hot-set class | 9 of 9 live hot sets refused (minimum sites 0.9809 to 0.9919), both single-item programs refused | under 0.33 | adv-accept, "the floor's final tally", 01:13 BST 8 October; class-v5 section 14 |
| (c'') floor at 0.98 on the low-entropy band | 5 of 5 (p23, p18, p19, p15, p56) | under 0.60 | attack-pass-2026-10.md, sub-version 3 |
| (a') freshness on the lineage-visible constant class | 11 of 11 (sub-version 1's eleven seeds under 1.2x on sub-version 3) | under 0.27 | attack-pass-2026-10.md, the sub-version 1 table |
| The bucket bound (in sigma of the bucket's expectation at 2^20) on the quarter-bit class | 4 of 4 would be caught (p4, p8, p10, p34: largest bucket 3.1x to 5.6x at narrow windows, +17 to +37 sigma), not yet wired as a rule | under 0.75 once wired; the clean spread is measured in section 6 | the hash lane's attribution, 09:46 BST 8 October |
| The index-bit bias read on the product class | 3 of 3 named sites (Devnet 3 site 0 at P(bit 0) = 0.25; era-drawn-28 site 15 at 0.778; the per-load candidate 0 at z 29.5), and 7 of 7 in the 16-era smoke run at 130 to 511 sigma | under 0.30 on ten cases; but the read is a record, not a refusal (section 3.2) | adv-cache-2 2.3; counter-asic-4-research.md 20.2a; section 6 |
| The day-key cost rule | the planted mul1all and mulnaf days, day 29,337 | 0 of 2^24 days over 1.1x after the rule (a census, not a case count) | f4-weakday.md, 8ca66afa |
Reading: the per-test miss rates are bounded by HANDFULS of known cases, so the union bound today is a loose number (the sum exceeds 1). The way to tighten it is not more eras but more known-failed cases per test, which the adversarial sweeps supply cheaply: adv-accept's tail holds 22 programs beyond the 1.2x gate of 37 measured live, and the drawn-era census of adv-cache-2 holds 19 biased sites of 1,056. A test with 60 fired cases has `delta` under 0.05; with 300, under 0.01. The proof of testing (section 4.3) therefore carries, per test, the count of known-failed cases it fired on, and that count is the number a reader checks first.
### 2.3 Stratified draws over the corners
The axes with a known mechanism and their corner values: `m` in {4, 16} (margin, verifier), weights at the lossy cap (or, mul, mulhi at +B) and at the table, `W` in {1, 4}, shape in {64, 256}, R in {1..27} against {28..31} (the product-bias class visible or hidden; at D = 29 the hidden set is {30, 31}), the windows forced to quarter (`k_off = 2` at every site, the F8 tail class), D in {28, 29}. That is 2 x 2 x 2 x 2 x 2 x 2 x 2 = 128 corner cells for the per-era statistics (the mixer cells are per-family rows), plus the random stratum (every axis drawn). The plan: at least 8 eras per cell for the cheap ring-B statistics (1,024 eras, about 2 core-hours each at the (c''') cost, 2,000 core-hours, 64 box-hours of 32-core slots) and the random stratum at 3,000 eras (the 2^-10 line of 2.1 at 95 percent for the ring-B tests, 6,000 core-hours, 190 box-hours); the live-dataset census at 2^24 on 64 seeds for 24 corner representatives and 64 random eras (88 eras at 45 to 95 core-hours each, about 6,000 core-hours, 190 box-hours). About 450 box-hours in all, 2 to 3 days of both boxes at the 88-core pool, inside the testnet period by a wide margin. Section 6 records what has run.
## 3. The per-era automatic tests (layer 4)
The three rings. A test's ring says who runs it and what a failure costs.
| Ring | Who runs it, when | Cost per run | What a failure does |
|---|---|---|---|
| A: structural, per era | every node at the era cut, on the drawn parameters | microseconds | the era stream's next block of draws is consumed (a redraw), up to a cap, then the base table (the class v4 values) as the last resort, so the draw is total and no consensus path panics (the AP-F8-2 shape) |
| B: the acceptance, per epoch | every node at the epoch draw, on each candidate | about 2.2 s per chosen candidate on one box core at 2^20 (the (c'') pass), plus 0.3 s for the 64-unit parts; the two next-class tests ride the same 2^20 pass | the next attempt, as today; the cap 256 and the last resort unchanged |
| C: offline, per drawn era | the family gate on the boxes, before the testnet and through it | box-hours | a finding against the family's band (a band narrowed, a rule added), never a chain event |
### 3.1 Ring A: the structural rules the chain runs per era
| Rule | Draw it reads | Known-failed shape | Clean refuse rate | Source |
|---|---|---|---|---|
| Band membership: `m` in the band, `W` in the allowed set, the shape in the set, every weight within B of the table with shfl and mulhi at or under their base, the sum 75 | the layer-1 draws | a draw outside the band (a unit test) | 0 by construction (the draw is bounded) | spec 1.13.1's form |
| The day-key cost rule (AP-F4-1, class v5) | ROT, MUL | mul1all, mulnaf, day 29,337 | 6.1e-4 per day | f4-weakday.md 6.7 to 8; 8ca66afa |
| The rotation class recorded: whether R leaves a product's low bits inside the index at the era's D | R, D | none (a record, not a refusal: restricting R to {29, 30} would cut the interleave's entropy from 31 to 2 values; the value-level test of ring B is the catch) | 0 | adv-cache-2 2.3 |
| The weights' lossy share bound: or + mul + mulhi at most the table's 18 plus B, so the exhaustion bound `r^256` stays under 1e-18 (r under 0.85) | weights | the lossy corner of section 6 if its `r` reads over 0.85 | 0 by construction | row 90 (r = 0.681 at the table) and section 6 |
### 3.2 Ring B: the acceptance rule generalised to the family, per epoch
| Part | Generalisation | Known-failed shape (fires today) | Clean refuse rate (the cost) | Per-candidate cost |
|---|---|---|---|---|
| (a) (b) | unchanged | the class v2 census vectors (22, 37, 51) | 11.6 and 3.0 percent of rejections | microseconds |
| (a') the freshness fixpoint over base then shadow | keyed on the family's shape (`is_family_shape`: shadow in {64, 128, 256} at 6,912 per iteration, `m` in {4, 8, 16}, the mix one-hot on the drawn width) in place of the one class v4 shape, in the draw's source rule and the acceptance alike | sub-version 1's eleven seeds | 83.5 percent of rejections (57.5 percent of candidates) | microseconds |
| (c) the dynamic test with the shadow executed | the interpreter already runs any block length, any pass count and any width (the fold) | p23 at attempt 1 (the shared-operand idiom) | 0.9 percent | about 3 ms (64 units) |
| (c') saturated sources | unchanged | the forced-saturation plant | 0.05 percent | in (c) |
| (c'') the distinct-index ratio at 0.98 | `E_s = N - N^2 / (2 W_s / W)`: the window divided by the width | p15, p18, p19, p56 | 1.1 percent | 2.2 s at 2^20 |
| (c''') the hot-item floor at 0.995 | the same `E_s`; the floor's clean spread re-read per stratum (section 6): a fixed 0.995 under a spread that moves with the shape or the width is either a liveness cost or a miss | seed 100767 (0.9919) and the nine live hot sets | 2.435 percent of accepted class v4 programs, +0.045 on the mean attempt | in the (c'') pass |
| The per-site largest-256-item-bucket bound (NEW, the F8 tail's catch) | the largest count over a site's 4,096-word buckets in its window against `N / (W_s / 4096)`, in SIGMA of the bucket's expectation (`(max - E) / sqrt(E)`), refused above a band set from the clean spread. The ratio alone cannot carry the bound: a clean full-window site's largest of 65,536 Poisson(16) buckets reads 2.1x at +4.4 sigma, a quarter-window site's largest of 16,384 Poisson(64) buckets 1.5x at +4 sigma (the 16-era smoke run read ratios 2.2 to 2.4 on clean sites); the four tail seeds' 3.1x to 5.6x at narrow windows are +17 to +37 sigma | p4, p8, p10, p34 | measured in section 6 (the columns `bucket_z_max`, `bucket_win`) | one linear pass over the (c'') indices |
| The index-bit bias read (NEW, adv-cache-2's value-level test) | per site, the one-count of each free index bit (bits above the width's alignment and below the window's cut) over the 2^20 evaluations, in sigma (`sigma = sqrt(N) / 2 = 512`); a 6-sigma band is a bias of 0.3 percent at 2^20 (the product law's bit 0 at 25 percent reads z about 512, bit 1 at 3/8 z 256, bit 3 at 3.1 percent z 64, bit 6 at 0.4 percent z 8, bit 7 at 0.2 percent z 4: passes). MEASURED on the first 16 drawn eras (11:21 BST, section 6): 7 of 16 accepted programs carry one site with a bias of 130 to 511 sigma at address bit R or R + 1 (the era's rotation), the era-stride class exactly, on programs (c''') passes at 0.9954 to 1.0000; the 9 others read under 3.8 sigma. So as a REFUSAL the read would redraw about 40 percent of epochs; its right use is as a record per era plus the structural lever (fold the product's low bits in `load_index` before the rotation, a next-class item), and the chip price is stated in section 5 | Devnet 3 site 0, era-drawn-28 site 15, the per-load candidate 0; and now the 7 of 16 | 7 of 16 at 6 sigma (a refusal is not the form); 0 of 16 at a 600-sigma band | one pass over the indices, 28 bit tests each |
| The cap and the last resort | unchanged at 256; the last resort's own verdict is adv-accept-3's finding 1 (9 percent of last-resort programs fail (a)); the family keeps the open fix (continue past 256 until one passes) | the reject-everything mirror | 0 reached in 1.02 x 10^6 seeds | n/a |
### 3.3 Ring C: the offline tests per drawn era
| Test | What it reads | Known-failed shape | Clean cost per era | Sample needed for its own verdict |
|---|---|---|---|---|
| The attempts census (row 90's form) | every candidate of each seed's chain draw through the real rule; `r` per candidate, the parts' shares, the attempt histogram, `P(exhaust) = r^256`, 0 at the cap | the forced-exhaust plant (hits the cap, prints the last resort); the per-load class (r about 0.98) | 1,000 seeds: 202 s on 23 threads at the table (1.3 core-hours); about 5 core-seconds per accepted program with (c''') | 1,000 seeds per stratum reads `r` to 0.01 |
| The F8-form uniformity census at 2^24 on the live dataset | the item histogram against the window model (the program's own 16 draws under the era's windows at the era's D), the top 0.1 percent within 1.2x, the hot-set test `X_f >= f`, the 6-sigma largest bucket, the per-site attribution | quarter-lines, half-lines, const-item (the plants); p31 (29.27x), p23 (4.82x) | 42 to 95 core-minutes per seed at 2^24 (the two halves of 32 on build-2 took 41 to 46 minutes at 30 to 32 cores; the 64 seeds of sub-version 3 took 69 minutes of the box); 64 seeds about 45 to 100 core-hours | 64 seeds per era bound the fraction of epochs over 1.2x at 3/64 (4.7 percent) at 95 percent: a calibration of ring B, not a hot-set hunt (adv-accept's sweep found 9 hot sets in 796,042 draws, 1e-5, every one refused by the floor) |
| The stand-in gap at the era's D | the (c'') and bucket statistics on the closed form at D = 28 against the live dataset at the era's D on the same programs | none yet (the gap read 0.0004 at D = 28; a D where the windows disagree is the plant to build) | one 2^24 run per program, about 1 core-hour; 54 programs per era class | per D value, not per era |
| The exhaustion count at scale | 10^4 to 10^5 chain-shaped seeds per stratum | the per-load class | 1.45 s per seed (F9's chain path): 4 core-hours per 10^4 | 10^5 per lossy corner reads the past-31 tail |
## 4. The board over the testnet period
### 4.1 F1 to F10 and the nine lanes, classified
| Row | Harness input | Class | Per what | Cost per run | Note |
|---|---|---|---|---|---|
| F1 shadow redundancy | the shape (block length, passes) | per family | each of the three shapes | 10^4 programs at 5 core-seconds each (the (c''') draw): 14 core-hours; 10^5: 140 | the 3.0 percent `SHADOW_REMOVABLE_MAX_PERMILLE` re-read per shape |
| F2, adv-mixer, adv-mixer-3 (the round margin) | `m`, the day | per family | `m` in {4, 16} on 10 days each | index census 2^32 t about 7 core-hours per k per day; sac at 2^24 similar; SAT 1 core-hour per k | at `m = 4` the gate "no distinguisher beyond 2 of m" leaves 2 of 4 |
| F3, adv-cache, adv-cache-3 (the chained cache) | the cache log2 (layer 2's doublings) | per family | each cache size step | minutes (the closure search), the curve by arithmetic | the fill is not drawn |
| F4, adv-mixer-2 (the weak day) | `m`, the day key | per family and chain-run per day | 2^24 days per `m` | 379 s on 12 cores (1.3 core-hours) | the rule is chain-run; the census is its calibration |
| F5 chip model | the band's top (`m = 16`, the dearest mix, `W = 4`, rung 2) | per family | once per band | minutes | the premium row per corner |
| F6 verifier worst case | `m`, shape, `W`, N, the fill at D | per family | the top corner (`m = 16`, rung 2, `W = 4`) | 10^5 programs timed on the half-core proxy: about 100 core-hours | the x16 row decides whether 16 enters the band |
| F7 era census | the era draw with the layer-1 draws added | per family | 2^24 era seeds | minutes | stride bijective, R, pos, M, the weights' corners uniform |
| F8 uniformity | the era (every axis) | per era (ring C) and chain-run per epoch (ring B) | 64 seeds per drawn era | 45 to 100 core-hours | section 3.3 |
| F9 acceptance edges, grinding, exhaustion | the era | per era (the exhaustion count) and per family (the 5090 grinding row) | 10^4 seeds per stratum | 4 core-hours | the GPU row once per band |
| F10 ladder monotonicity | N | per family | once | the fast-time harness, minutes | unchanged by the draws |
| adv-cache-2 (hot set, era stride) | the era | per era | 32 drawn programs at 2^23 | 0.7 core-hours per program | the value-level read is ring B now |
| adv-accept, adv-accept-2, adv-accept-3 | the era | per era (the sweep and the attempts census) and per family (the A6000 header row) | as the attempts census | as above | the last-resort fix is the family's |
### 4.2 The testnet period as the window
Before the first miner: the per-family rows at the corners (F1 at three shapes, the mixer ladder at `m = 4` and 16, F4 per `m`, F6 at the top corner, F7 with the new draws), and the ring-C census on the corner cells and the random stratum of section 2.3. Through the period: every era the chain actually draws is a drawn era of the census (the chain's own eras are the best sample, the one the rule of three speaks about), so the F8-form census at 2^24 runs on each live era within its first day, and the attempts census on each; a live era that fails ring C is a finding against the band, filed and priced, never a chain event (the chain's ring-A and ring-B tests already passed it). The board's verdict at the end of the period: the count of live eras drawn and passed, the corner cells drawn and passed, the worst reading per test, the refuse rate per stratum, and the bound's arithmetic from those counts.
### 4.3 The proof of testing (the published format)
One JSON record per drawn era, in `docs/analysis/class-v6/gate-records/<family-id>/<stratum>/<era-label>.json`, the shape of `docs/plans/counter-asic-3-gate/*.json`:
```
{ "family": { "id": "<sha256 of the genesis band table>", "spec": "docs/spec/01-lottery-hash.md@<commit>", "object": "igneum-pow@<commit>" },
"era": { "label": "igneum-family-gate/era/<k>", "stratum": "random | lossy-cap | w4 | shape64 | ...",
"draw": { "m": 8, "weights": [12,10,8,8,8,7,6,6,6,4], "width_words": 1, "shape": [256, 27], "R": 17, "M": "0x9ad30d99", "pos": [0,2,12,13] } },
"harness": { "binary": "<sha256 of the pinned copy>", "box": "igneum-build-1", "lease": "<the lease line>", "started": "<UTC>", "ended": "<UTC>", "core_seconds": 7210 },
"tests": [ { "name": "attempts-census", "seeds": "igneum-family-gate/program/<k>/epoch/0..999", "r": 0.681, "parts": {"a_prime": 0.835, ...}, "max_attempt": 28, "exhausted": 0, "verdict": "PASS", "log": "<path>", "log_sha256": "<hex>" },
{ "name": "c3-floor", "floor": 0.995, "min_ratio": 0.9962, "min_site": 6, "verdict": "PASS", ... },
{ "name": "bucket-bound", "bound": 2.0, "max_ratio": 1.31, "verdict": "PASS", ... },
{ "name": "bit-bias", "band_sigma": 6, "max_z": 2.4, "verdict": "PASS", ... },
{ "name": "f8-census-2e24", "seeds": 64, "over_1.2x": 2, "worst": 1.38, "hot_sets": 0, "verdict": "PASS", "known_failed_fired": ["quarter-lines", "const-item"], ... } ],
"label": "internal research, not an independent review" }
```
And one family summary per board cut, `docs/analysis/class-v6/gate-records/<family-id>/summary.md`: eras drawn per stratum; per test the count passed, the worst reading, the refuse rate, the known-failed cases fired (the count that bounds `delta`); the rule-of-three line per test (`3 / n` at 95 percent); the union bound's sum; the box-hours; the list of corner cells not reached. A reader checks the record by re-running one era from its label with the pinned binary; every number in the summary is a count over the records.
## 5. The honest line
What the family gate cannot see, and how a chip would use it.
| Blind spot | Why the gate misses it | What a chip gets from it | Bound today |
|---|---|---|---|
| The era-stride class, measured today at the bit level on 7 of 16 drawn eras (section 3.2): not diffuse, a 25 percent bias on ONE address bit (bit R, the product's bit 0 through `rotl(x * M, R)`) at one site, on programs every floor passes | the (c''') floor reads distinct counts, which a one-bit bias barely moves (0.9954 to 1.0000 on the seven); the bit read sees it at 130 to 511 sigma but refusing it would redraw about 40 percent of epochs; the class is structural (`load_index`'s form and the weights' products), so the catch is a next-class change to the address form, not a per-era floor | a store holding the favoured half of that site's window (128 MiB at D = 28) serves 75 percent of that site's reads instead of 50: 25 percent of 1/16 of a hash's reads, about 1.6 percent of reads at f = 1/2 from one site in about 40 percent of eras (arithmetic on the measured bias); the honest partial-store curve already costs 1.26x the ops at f = 1/2 (adv-cache-2 Q4), so the f = 1 verdict stands; the top-0.1-percent reading of the ledger (1.0024x) was the same class seen through the coarser instrument | the per-era record (ring A names the rotation class; ring B logs the bit read); the structural fix is the family's first owed item for the research lane |
| A joint corner no stratum names | the 128 cells pin the axes with a KNOWN mechanism; a mechanism that needs, say, a particular `M` with a particular weight table and a particular shape is reached only by the random stratum, which bounds its fraction at 3/n | one era (180 days) of whatever the corner is worth, at most once per occurrence | 1/64 to 1/256 of eras untested at 95 percent (section 2.1); the chain-run rings catch the known shapes on every era |
| The mixer below adv-mixer-3's bands | biases under 0.00018 (Q2b), correlations under 0.00073 with multi-bit masks, differentials under 2^-19, any structure a SAT model of 2 or more applications would expose; at `m = 4` the measured margin is 2 applications | a shortcut that saves applications against the 9,360 ops per item; none found at any k | the round margin, stated; `m = 4` is the corner that spends 4 of the 6 measured spare applications |
| The verifier at `m = 16` and rung 2 on a 2019-class core | unmeasured (an estimate of 9 ms against the 10 ms gate) | nothing for a chip; a node tier retired if the corner is drawn | the row is owed before 16 enters the band |
| The acceptance's stand-in at D = 28 against a live D = 29 | the windows cap `min(k_off, D - 26)` and the rotation's cut set move with D; the gap is measured at D = 28 only | a per-site concentration the closed form does not reproduce | 0.0004 at D = 28; unmeasured at 29 (ring C row) |
| The union bound's own looseness | the per-test miss rates rest on 3 to 11 known-failed cases each (section 2.2), so the family's false-pass rate is not yet a number under 1 | nothing new; the bound is on what the gate claims, not on the hash | fixed by cases, not eras: 60 fired cases per test for `delta` under 0.05 |
| The last-resort program | 9 percent of last-resort programs fail (a) and are handed out unchecked; a lossy-corner era that raised `r` toward 0.95 would reach it at 2e-6 per epoch | a program that re-reads one address in two loads of the same hash | ring A's lossy-share bound keeps `r` under 0.85 (`r^256` under 1e-18); the fix (continue past 256) is owed to the rule's owner |
| The era redraw's own grindability | a ring-A redraw consumes the stream's next block, so an adversary who could steer `E_n` could steer which block is used; the era VDF of F7 (517 s on the fastest prover, 259x the window) closes the steering of `E_n` itself | nothing beyond F7's bound | F7 (a) PASS with the VDF in the node |
A GPU-like chip (the `f = 1` stored-dataset chip with a programmable core, chip-model-v3 section 5) loses nothing on any era of the family: every drawn parameter is firmware to it, and what the family costs it is the core sized for the band's top (the research lane's USD 30 to 60 per chip, modelled). The family gate's job is therefore not to defeat that chip (the identity of the research file's section 2 says the per-joule edge is the card's own idle, 2.1x at `k = 1`) but to make sure no drawn era hands ANY chip more than the one hash did: a per-era hot set, a per-era mixer shortcut, a per-era verifier miss. Everything in this document is a bound on that increment, and the increments found so far are 1.002x and 1.0024x.
## 6. Running now, and the coverage measured
Started 11:1x BST, 8 October 2026, on igneum-build-1 under `lease pool` (class measure, owner class-v6-family-gate), the scripts under `/srv/builds/_adv-family-gate/` (fg-setup.sh, fg-build.sh, fg-census.sh), every binary run from a pinned copy under `bin/<tag>/` with its sha256 (AP-H2's rule), logs and TSV rows under `logs/`, copies of the finished logs land under `docs/analysis/class-v6/logs/` on this branch with the full report.
| Run | What | Where | State |
|---|---|---|---|
| base | `v5_attempts_census` (the shipped class v5 draw: shape 256 x 27, `m = 8`, `W = 1`, the table weights) over f8-label seeds 1,000 to 11,000, each seed its own drawn era (M, R, pos, windows) | build-1, 24 cores, binary 4bec799c (8f481459 unmodified) | running from 11:16 BST |
| fg1 random | `family_gate_era_census` with every axis drawn (shape in {64, 128, 256}, `m` in {4, 8, 16} recorded, `W` over {1, 4}, weights at B = 4 with shfl and mulhi never raised) | build-1 | the harness built at 11:2x BST; the smoke run and the chunks follow |
| fg1 corners | the same harness pinned per cell: `IGNEUM_FG_LOSSY_CAP`, `IGNEUM_FG_WIDTH=4`, `IGNEUM_FG_SHAPE=64`, the combinations | build-1, build-3, build-4 as the pool frees | queued |
The coverage table (filled from the TSV rows at 16:xx BST for the 17:00 cut; the full read at 09:00 BST tomorrow):
| Stratum | Eras drawn | `r` per candidate | Mean attempt, max | Exhausted | min (c'') ratio (the worst era) | Under 0.995 (the (c''') refuse rate on the accepted draw) | Largest bucket ratio (worst era, site) | Largest bit bias in sigma (worst era, site, bit) |
|---|---|---|---|---|---|---|---|---|
| base (shipped draw) | pending | | | | | | | |
| random (every axis drawn) | pending | | | | | | | |
| lossy cap | pending | | | | | | | |
| W = 4 | pending | | | | | | | |
| shape 64 | pending | | | | | | | |
## 7. Owed by 09:00 BST on 9 October (the full report)
1. The coverage table of section 6 filled from the logs, with the per-stratum refuse rates and the union-bound arithmetic on the measured numbers, and the TSV rows and logs under `docs/analysis/class-v6/logs/`.
2. The clean spread of the bucket ratio and the bit bias per stratum, so the two new ring-B tests get a bound and a clean refuse rate (not a guess).
3. The lossy-corner `r` against the 0.85 line, and the exhaustion count at 10^4 seeds on that corner.
4. The width-4 reading of (c'') and (c''') with the divided expectation, and the spec line for 1.4.6.5.
5. The mixer rows at `m = 4` and 16 (adv-mixer-3's index and sac at k = 2 on two days; the SAT at k = 2) as the per-family corner rows, if the boxes have the cores after the other lanes' jobs arrive; else the plan with its hours.
6. The F8-form census at 2^24 on 64 seeds of one lossy-cap era and one width-4 era, the first ring-C rows on the family (about 2 box-hours each).
7. The gate-record JSON written for every era of the census from the TSV (the format of 4.3), and the summary.
## 8. Sources
In-tree (the mirror's master and lane branches at 8 October 2026, 11:00 BST): `docs/plans/cryptanalysis/in-house-pass.md` (sections 13 and 14, the nine lanes' close); `docs/analysis/attack-pass-2026-10.md` (F1 to F10; the sub-version 1, 2 and 3 re-gate tables; the tail attribution of 8 October, 09:40 to 09:46 BST); `docs/analysis/cryptanalysis/report-mixer-3.md` on branch adv-mixer-3 at 981bfff2 (the round margin); `docs/analysis/cryptanalysis/report-chained-cache-2.md` on adv-cache-2 at bfc3746c (section 2.3, the era-stride table); `docs/analysis/cryptanalysis/report-acceptance-rule-3.md` on adv-accept-3 at 7826d2b2 (Q1, Q1b); `docs/analysis/cryptanalysis/report-acceptance-rule.md` on adv-accept at 8f188e5a (row 90, the floor's tally, gap-deep); `docs/spec/01-lottery-hash.md` 1.4.6, 1.4.7, 1.8.4, 1.13; `docs/design/class-v5-stored-state.md` on class-v5 at 8f481459 (sections 11 and 14) and `docs/design/class-v5-harness/` (v5-attempts-census-1000.log, v5-census-4600-0.log); `docs/plans/counter-asic-3-status.md` section 7c; `docs/plans/read-width.md`; `docs/analysis/counter-asic-4-research.md` on counter-asic-4 at 7a133d76 (15.1a, 15.1b, 20.2a, 20.2b, 20.3); `docs/design/class-v6-rotating-family.md` on the same branch (the outline).
Public, read 8 October 2026:
- Hanley, J. A. and Lippman-Hand, A., "If nothing goes wrong, is everything all right? Interpreting zero numerators", JAMA 249(13), 1743 to 1745, 1983; the rule of three as a one-sided 95 percent bound 3/n on a binomial rate with zero events in n trials: https://en.wikipedia.org/wiki/Rule_of_three_(statistics)
- Wilks, S. S., "Determination of Sample Sizes for Setting Tolerance Limits", Annals of Mathematical Statistics 12(1), 91 to 96, 1941; the nonparametric tolerance limit on the extreme order statistic, the same arithmetic as 2.1: https://www.projecteuclid.org/euclid.aoms/1177700380
- Böhme, M., "STADS: Software Testing as Species Discovery", ACM TOSEM 27(2), 2018; the Good-Turing estimate `f_1 / n` of the discovery probability as the residual-risk bound of a testing campaign with no finding (the fuzzing-coverage view of section 2): https://arxiv.org/abs/1803.02130
- OSTIF, "Four audits of RandomX for Monero and Arweave have been completed: results" (2019; Trail of Bits, X41 D-SEC, Kudelski Security, Quarkslab; no critical finding; the one-round AES diffusion and the BLAKE2 rows): https://ostif.org/four-audits-of-randomx-for-monero-and-arweave-have-been-completed-results/
- Quarkslab, "Security audit of Monero RandomX" (2019; the "alternative configurations" caveat: the audits bounded the shipped configuration, which is the single-member reading this document generalises): https://blog.quarkslab.com/security-audit-of-monero-randomx.html
- Least Authority, ProgPoW algorithm audit (final report 9 September 2019; the Ethereum Cat Herders, the Ethereum Foundation and Bitfly as clients): https://leastauthority.com/blog/2019/09/09/
- The hardware audit of ProgPoW (Bob Rao, 2019; "works well against conventional ASIC strategies", the memory-intensive threat kept open), as reported: https://criptonoticias.com/mineria/desarrolladores-ethereum-presentan-resultados-auditoria-progpow

View file

@ -0,0 +1,443 @@
# Class v6 research lane B: the hardware future, five years out
Lane B of the class v6 rotating-family research (the founder's word of 8 October 2026, 11:1x UK: "see if anything can be
optimised, added or invented"). First cut landed 8 October 2026, 11:5x UK; the full report fills the same file. Every
figure carries a label: **measured** (a number read off an instrument in this repository, with the file), **claimed**
(a vendor's or a paper's number, with the URL and the date read), **modelled** (arithmetic on claimed figures by the
method of `docs/analysis/chip-model-v3.md` section 5), **approximate** (from memory or an estimate; the sensitivity is
given). Nothing here is a measurement of a chip. Reading public research is in-house; nothing was paid for or asked of
anyone outside.
The question, as the coordinator put it: for each memory or packaging line, what does it do to a chip's cost per
dependent random read over a dataset of 1 to 8 GiB that grows with chain state (energy per read, latency, capacity cost
per GB, availability to a non-hyperscaler), what k band does it give the chip five years out, and which ONE of the four
v6 layers (1: per-era parameter draws; 2: the state-sized dataset with a floor; 3: scheduled family epochs; 4: the
(c''') acceptance floor and the F8 uniformity test per era) blunts it, with a number. Where a line beats every layer,
this file says so with the number.
## 0. One page
The reads are the hash. Under class v3 the RTX 5090 spends 2.40 microjoules per hash on 128 dependent reads, 18.8 nJ
per read all-in, and the memory system itself spends 2.0 nJ of that (chip-model-v3 5.3, modelled); the card's own
marginal per dependent DRAM read is 10.9 nJ unlocked and 8.7 nJ at the 1,300 MHz lock (measured 8 October 2026,
counter-asic-4-research 15.1a). Every chip in this file is a machine that pays the memory's nanojoule and not the
card's ten, plus whatever the shadow (the program work drawn into the memory wait) forces it to pay at `k` times the
GPU's cost per op. That identity does not change with any technology below; what changes is the memory's nanojoule,
the rate a chip can read at, what a GB costs, and who can buy it.
**The three findings that change v6's design**
1. **The strongest five-year chip is not a DRAM chip. It is a 2 GiB SRAM full store on one reticle of merchant N2, and
none of the four layers reaches it.** TSMC N2 reads 38 Mb/mm^2 of SRAM (claimed, IEEE Spectrum, 12 December 2024,
volume in 2025), so 2 GiB is about 452 mm^2 of macro, one die under the 858 mm^2 reticle, roughly USD 400 to 600 of
silicon at a USD 30,000 wafer (approximate). It has no activate ceiling, so its rate is power-bound: about 2,100
MH/s at 300 W, 0.14 microjoules per hash, **13x to 17x the 5090 per joule at zero shadow and about USD 0.3 per
MH/s** against the 5090's 14.7 and the GDDR7 chip's 2.8 (modelled; the wire energy, 0.5 to 2.0 nJ per read, is the
sensitivity). Layer 2's floor moves its capex, not its joules: at 4 GiB it is two dies, at 8 GiB four, USD 1,000 to
2,500 of silicon, and its edge per joule falls only from 17x to about 13x because an inter-die hop costs 0.27 nJ
(UCIe, claimed 0.5 pJ/bit). A floor that would blunt it on joules (16 to 32 GiB, eight to sixteen reticles) retires
every honest card under 32 GB first. The only lever that reaches it is the shadow at `k`: with the shadow core on
the same N2 die the ALU band is 0.3 to 0.8 (the record's), and the class v4 premium holds **4.8x at k = 0.5, 2.7x
at k = 1**; at the 5090's whole ALU budget (about 330,000 ops per hash, 575 W) 3.7x and 2.0x. The design change:
layer 1's program-length draw is sized against this chip, not the GDDR7 board, with its lower bound at the shadow
that holds the record's 2.1x today and its upper bound at the honest cards' full latency shadow, re-based every
family epoch (layer 3) on the cards then mining; and the public "2x" line is not reachable in this model against
this chip at any `k` under 1. What slows it is money and time (an N2 project, USD 100 M to 500 M and 18 to 24
months, approximate), which is the clock of Counter ASIC 3.0 item 4, not a hash property.
2. **Per-bank processing-in-memory is structurally blind to this hash; the real near-memory threat is the custom
HBM4E base die, and layers 1 to 4 do nothing to it.** HBM-PIM, AiM, LPDDR5X-PIM and UPMEM put a compute unit beside
each bank (or a DPU per 64 MB); a dependent read's next address is uniform over the dataset, so it lands in the same
bank with probability bank bytes over dataset bytes: **1.6 percent at 2 GiB on a 32 MB bank, 0.4 percent at 8
GiB**, and UPMEM has "no direct communication channel among DPUs" (claimed, the PrIM paper), so the other 98
percent of reads go to the host. The unit that can follow the chain across banks is a controller on the stack's
base die, which is exactly what TSMC's custom C-HBM4E is: "the custom base die will integrate memory controllers and
PHY" on N3P at 0.75 V, "2x the power efficiency", Micron production 2027, SK hynix HBM4E in 2026 (claimed,
TrendForce, 1 December 2025). That is the `f = 1` chip of the record with its controller moved into the stack: about
0.9 to 1.0 nJ per read, and HBM4's 32 channels (JEDEC JESD270-4, April 2025) double the activate-bound ceiling per
stack, so **6.5x to 14x per joule at zero shadow** (the ceiling unmeasured, as the record's HBM rows are). The four
layers act on the program, the item map and the capacity; this chip runs any program and holds 36 to 64 GB. What
brakes it until about 2028 is availability (HBM allocated to AI, 20 to 26 week leads, Samsung asking USD 4 to 5 per
Gbit for HBM4 against 1.5 for HBM3E, October 2026) and the shadow at `k`: 4.4x at k = 0.5, 2.6x at k = 1.
3. **The denominator is the wrong card.** Every chip edge in the record is quoted against the 5090 at 2.40 microjoules.
The Apple M5 Max measured 0.78 microjoules per hash at the GPU-plus-DRAM meter (latency-shadow-2026-10-06, 6
October 2026), 3.1x the 5090 per joule, and LPDDR6 SoCs (JEDEC JESD209-6, 2025; 14.4 Gbps, 32-byte atoms) are that
tier's next step. Against the M5 Max the GDDR7 chip reads 1.7x, one HBM3 stack 2.4x, the HBM4 base die 3.6x to
4.4x, the N2 SRAM die about 5x, and with the shadow at k = 0.5 the SRAM die reads about 3x. The honest joule, not
the 5090's, is the chain's resistance, and the same chips are 2x to 3x less frightening against it. The design
change: v6's acceptance floor (layer 4) and the shadow sizing (layer 1) are scored per card tier with the
unified-memory SoC tier as the reference joule, the dataset is kept inside 16 GB unified memory (8 GiB at the top of
the schedule does that), and the public text states the edge over the best honest joule, which is the number a
miner can act on.
The honest line, in one sentence: five years out a 2 to 8 GiB dataset fits in one to four reticles of merchant SRAM at
USD 500 to 2,500 of silicon, every DRAM line converges on the same 0.5 to 1.0 nJ per read with its controller in the
stack, and no layer of the four touches either; the shadow at the measured `k` band holds 3x to 5x, the honest tier's
own efficiency halves that again, and the clock (project cost against daily issuance) is the wall that is left.
## 1. The method and the denominators
| Quantity | Value | Label | Source |
|---|---|---|---|
| The 5090 at class v3 | 136.1 MH/s at 326 W, 2.40 microjoules per hash, 17.5 G dependent reads per second, 415 ns at 256 lanes, a 32-byte sector per 4-byte read | measured | `docs/bench-log.md`; chip-model-v3 5.1 |
| The 5090 per dependent DRAM read, the whole card's marginal | 10.9 nJ unlocked, 8.7 nJ at the 1,300 MHz lock (dram_chase_1g, 18.17 G reads per second) | measured, 8 October 2026 | counter-asic-4-research 15.1a |
| The 5090 L2 hit, per dependent read | 2.4 nJ unlocked, 1.4 nJ at the lock | measured | the same |
| The 5090 per counted ALU op (the class v4 shadow's mix) | 11.3 pJ unlocked, 6.2 pJ at the lock (int_arx); the shadow's own 10.4 and 6.5 | measured | the same; counter-asic-4-research section 1 |
| The 5090 at the 1,300 MHz lock, class v3 | 134.6 MH/s at 223 W, 1.67 microjoules | measured, 7 October 2026 | counter-asic-3-status 7c, the clock grid |
| The class v4 premium `F` on the 5090 | 1.10 microjoules unlocked (145 W), 0.65 at the lock (82 to 88 W) | measured | the same |
| Apple M5 Max, class v3, GPU plus DRAM channels | 27.08 MH/s at 21.0 W, 0.78 microjoules; class v4 at 100,000 ops 1.40 (+16 W); 6.9 pJ per counted op | measured, 6 October 2026 (the SoC's other rails and the wall are not in the 21 W) | `docs/analysis/latency-shadow-2026-10-06.md` section 3 |
| RX 9070 XT, class v3 | about 18.6 MH/s at about 304 W, 10.6 microjoules; 2.4 G reads per second | rate measured, watts approximate | `docs/analysis/horizon/algorithm.md` 5.1 |
| The memory system's own cost per random 32-byte read | GDDR7 2.0 nJ, HBM3 1.2 nJ | modelled on O'Connor et al. 2017 Tables 2 and 3 and Samsung's pJ per bit roadmap | chip-model-v3 5.3 |
| The record's `f = 1` chips at zero shadow | GDDR7 board 166 MH/s at 77.6 W, 0.466 microjoules, 5.1x; one HBM3 stack 83.6 MH/s at 26.8 W, 0.321, 7.5x; eight stacks 666 MH/s at 174 W, 0.262, 9.2x | modelled | chip-model-v3 5.4 to 5.6 |
| The chip's `k` on the shadow's work | ALU 0.3 to 0.8; int8 tile 0.03 to 0.3 (the worse lever); L2 hit 0.1 to 0.3; shuffle 0.4 to 0.7 | the GPU side measured, the chip side claimed | counter-asic-4-research 15.1a |
The chip's energy per hash is `E = 128 x E_read + P_static / rate + k x F`, the rate the smaller of the memory's
activate-bound ceiling over 128 and the power budget over the per-hash energy, and the edge is the card's microjoules
over the chip's. Two `k` columns appear below: `k_read`, the memory's energy per dependent read over the 5090's measured
10.9 nJ (what the technology itself buys, independent of any program), and the shadow `k` band, the chip's cost per
forced op over the GPU's, which no memory technology changes (it is a logic-node question: a chip core at N2 against a
GPU at N3 or N4 sits at the band's low end, 0.3; at the same node, 0.5 to 0.8; approximate).
Reads are counted at the part's atom: 32 bytes on GDDR7, HBM3, HBM4 and LPDDR6 (each gives a 32-byte minimum access;
JEDEC, claimed), 64 bytes on an SRAM macro or a line-based part. The class v6 read-width draw (4 to 64 bytes per read
under layer 1) does not move any row: a chip's controller fetches the atom whatever the hash asks for, as the 5090
fetches its 32-byte sector for a 4-byte load, and the record's w64 reading (the 5090 bandwidth-bound at 71.9 MH/s,
chip-model-v3 5.7) says a wider honest read costs the card and gives the chip nothing. The dataset's size enters only
through capacity cost and die count; its growth (layer 2) enters through the number of dies or stacks a chip must carry
at each family epoch.
## 2. The table
Energy is per dependent random read at the part's atom. Latency is the read's own (controller to data), not the GPU's
queueing. Cost per GB is factory-gate where a source gives it, with contract pricing about 2x (Silicon Analysts, October
2026). "Edge" is per joule against the 5090 at class v3 and zero shadow; the bracket is against the M5 Max's 0.78.
"Year" is when a non-hyperscaler could put the part on a board. Every chip-side figure is modelled unless labelled.
| Technology | Year | Energy per dependent read | Latency | Cost per GB | Availability to a non-hyperscaler | `k_read` (over 10.9 nJ) | Edge at zero shadow, per joule | The layer that blunts it, and by how much |
|---|---|---|---|---|---|---|---|---|
| GDDR7 on a PCB, 16 devices, a 28 nm controller (the record's `f = 1` chip) | now | 2.0 nJ (modelled); 4.5 pJ per bit streaming (claimed, Micron) | about 55 ns controller, tRC about 45 ns | USD 10 (2 GB parts, ending) to 20 to 23 (3 GB parts at USD 60 to 70), September 2026 (claimed, TrendForce) | anyone, through distribution; the 2026 DRAM price cycle (LTAs USD 7.8 to 21 per GB, May 2026, claimed) roughly triples the chip's memory bill against 2025 | 0.18 | 5.1x (1.7x); 166 MH/s at 78 W | none of the four; the shadow at k = 1 holds 2.1x (the record); the dataset's size does not matter because the chip over-provisions capacity for channels (16 devices whatever the dataset) |
| GDDR7 at 36 to 48 Gbps, 4 and 6 GB devices | 2027 to 2028 (claimed, Micron roadmap via Guru3D and OC3D) | the same 2.0 nJ: the activate ceiling and the row energy do not move with the pin rate | the same | per-GB falls, per-channel rises: a random-read chip buys channels, not bytes | anyone | 0.18 | 5.1x | none; denser devices HURT the chip (fewer channels per GB), so this line goes the chain's way |
| HBM3E, one stack, a 28 nm controller on a one-stack interposer | now | 1.2 nJ (modelled); 4.05 pJ per bit streaming (claimed, Samsung) | about 50 ns | USD 8.3 factory gate (USD 300 per 36 GB), about 2x on contract (claimed, Silicon Analysts, October 2026); plus about USD 200 of interposer | Tier-1 volume pricing, 20 to 26 week leads, "supply constrained through 2026" (claimed); a one-stack buyer pays broker prices | 0.11 | 7.5x (2.4x); 84 MH/s at 27 W (ceiling unmeasured, 10.7 G reads per second; 2.3 G at the JEDEC tFAW, which would read 1.8x) | none; the shadow at k = 1 holds 2.4x (the record) |
| HBM4 (JESD270-4, April 2025): 2,048-bit, 32 channels x 2 pseudo-channels, 8 Gbps, 2 TB/s, 36 to 64 GB, VDDQ 0.7 to 0.9 V, a TSMC N12 base die at 0.8 V "1.5x efficiency" (claimed) | 2027 to 2028 for a non-hyperscaler | 1.0 to 1.1 nJ (modelled: the I/O term at 0.8 V against 1.1 V, the 909 pJ row activation unchanged) | about 50 ns | USD 15.3 (USD 550 per 36 GB, October 2026 estimate); Samsung asking USD 4 to 5 per Gbit for 2027 (USD 32 to 40 per GB) (claimed) | allocated to AI accelerators through 2027; "mid-to-high $4 per gigabit" in this month's negotiations (claimed, 2 October 2026) | 0.10 | 4.5x to 11x (1.5x to 3.6x): the 32 channels double the activate-bound ceiling per stack if tFAW is per channel (unmeasured) | none of the four; the shadow at k = 0.5 holds 4.4x, at k = 1 2.6x |
| Custom HBM4E base die (C-HBM4E): the controller and PHY in the stack on N3P at 0.75 V, "2x the power efficiency" (claimed, TrendForce, 1 December 2025); Micron 2027, SK hynix 2026 | 2028 or later for a non-hyperscaler | 0.9 to 1.0 nJ (modelled: no interposer crossing for the controller's traffic) | about 45 ns | HBM4E class, USD 20 to 40, plus a custom base die (USD 50 to 100 per stack, approximate) | by design a per-customer product of the three DRAM makers; the named customers are NVIDIA and Google; a miner-maker of Bitmain's size could commission one (approximate) | 0.09 | 6.5x to 14x (2.1x to 4.4x) | **beats all four layers**; the shadow at k = 0.5 holds 4.4x, at k = 1 2.6x; the shadow core sits on the same N3P base die |
| Fine-grained, hybrid-bonded DRAM on logic (FGDRAM-class 256-byte rows, "tFAW effectively eliminated", 51 percent lower energy per access, 4x bandwidth, claimed, O'Connor et al., MICRO 2017; hybrid-bonded stacks and 3D DRAM on the 2028 to 2030 roadmaps, approximate) | 2029 to 2031 | 0.5 to 0.7 nJ (modelled: 230 pJ activation for a 256-byte row plus about 1.2 pJ per bit of movement over bonded pads) | about 40 ns | HBM class, USD 15 to 40 (approximate) | custom, the three DRAM makers, top customers first | 0.05 | 12x to 20x (4x to 6x); about 125 MH/s per stack at 19 W, 1,000 MH/s at 150 W for eight | **beats all four layers**; the shadow at k = 0.5 holds 5.1x, at k = 1 2.8x |
| Per-bank PIM: Samsung HBM-PIM (a PCU per bank, 2021), SK hynix GDDR6-AiM (16 Gbps, 1.25 V, 2022) and AiMX (32 GB card), Samsung LPDDR5X-PIM (Hot Chips, 25 August 2026, 614 GB/s internal) and LPDDR6-PIM (JEDEC work "substantial progress") | now to 2027 | the host memory's own: 1.8 nJ (HBM2), 2.0 (GDDR6 class); the PIM unit never sees the read it would need | the same | the host part's plus a premium | HBM-PIM and AiM were samples and prototypes; LPDDR5X-PIM mass production "as early as 2027" (claimed) | n/a: no path for a cross-bank dependent read | none beyond the base-die row: the unit serves 1.6 percent of reads at 2 GiB (a 32 MB bank), 0.4 percent at 8 GiB | layer 2, trivially: any dataset over a bank; the number is bank bytes over dataset bytes |
| UPMEM DPU-in-DRAM: 128 DPUs per 8 GB DIMM, 64 MB MRAM per DPU, 350 MHz, MRAM by DMA at most 628 MB/s per DPU, "no direct communication channel among DPUs" (claimed, the PrIM paper, 2021); 1.2 W per 4 Gb chip, 10x DRAM's price at sampling, 1.5x projected (claimed, The Next Platform, 2020) | now | about 0.3 microseconds per 64-byte DMA at 350 MHz (approximate from the paper's alpha-plus-beta model); 2 GiB spans 32 DPUs and every cross-DPU hop is a host round trip of microseconds | microseconds | 1.5x to 10x DDR4 | anyone, in server DIMMs | n/a | none: dead at any dataset over 64 MB | layer 2's floor alone (2 GiB = 32 DPUs with no path between them) |
| LPDDR6 (JESD209-6, 2025): 24-bit channels as two 12-bit sub-channels, 32-byte minimum access, 10.7 to 14.4 Gbps, 4 to 64 Gb dies (claimed, JEDEC via PCWorld and HotHardware) | 2026 to 2028 | 1.5 to 2.0 nJ (approximate: short rows, low-voltage I/O over a PCB, 16 to 32 banks per channel) | about 60 ns | USD 8 to 21 in the 2026 cycle (LTAs, claimed); about USD 3 to 4 in 2024 (approximate) | anyone; it is also the honest SoC tier's memory | 0.15 to 0.18 | about 4x to 5x for a 32-channel controller chip (1.3x to 1.6x against the M5 Max, which already IS this part with a GPU) | none needed: the honest tier converges on it; see finding 3 |
| 3D DRAM (Samsung VS-DRAM 16 layers, VCT 4F2; Neo 3D X-DRAM 230 layers; SK hynix and Micron stacked cells), "from 2030" (claimed, heise, 29 May 2024) | 2030 or later | no vendor claims a latency or activate change; the row energy class is the planar part's (approximate) | the same | lower per GB after 2030 | the three makers | 0.1 to 0.2 | the DRAM rows above | none needed inside five years; it lowers the chip's and the card's capacity cost alike |
| RLDRAM 3 (Micron, 2011): tRC under 10 ns, 16 banks, no activate command, "SRAM-like random access" (claimed, Micron); 576 Mb and 1.125 Gb devices at 2,133 Mb/s (approximate) | now, in small volume | 2 to 4 nJ (approximate: a full small-row access per read; no energy figure found, the datasheet's power calculator exists) | under 10 ns | USD 400 to 800 (USD 50 to 100 per 1.125 Gb device, approximate, a networking part) | anyone, catalogue | 0.2 to 0.4 | about 2 G random reads per second PER DEVICE (16 banks over 8 ns), 32 G for 16 devices: 250 MH/s per board, about 5x per joule (approximate) | layer 2: USD 3,000 to 6,000 of memory at 8 GiB; the line to watch is AMD's "Folded Banks" (ISCA 2025): 8x the activate parallelism in HBM gives 6.7x the irregular bandwidth (claimed, the abstract) |
| FPGA with HBM2e: AMD Alveo V80 (Versal XCV80, 32 GB, 820 GB/s, 190 W, USD 9,495 MSRP, May 2024, claimed), Versal HBM VH1782, Altera Agilex 7 M (two HBM2e stacks, 820 GB/s), Alveo U55C USD 4,747 | now | HBM2 class, 1.8 nJ, but the fabric's controller reaches only the JEDEC tFAW ceiling: 2.3 to 2.4 G reads per second per stack (Shuhai, FCCM 2020; the record) | about 100 ns | USD 300 per GB of card | anyone, 10 to 24 week leads | 0.17 | 0.6x (0.2x): 4.8 G reads per second for two stacks, about 37 MH/s at 150 W, at 5x the 5090's price | none needed; no HBM3E FPGA was found announced (unverified) |
| Wafer-scale SRAM: Cerebras WSE-3, 44 GB SRAM, 21 PB/s, 900,000 cores, 46,225 mm^2 on N5 (claimed, Cerebras and The Next Platform, March 2024); CS-3 about 23 kW and "maybe $2.5 million" (claimed, approximate) | now, by the system | about 7 nJ (approximate): a 512-bit reply crossing about 140 mm of mesh on average at about 0.1 pJ per bit per mm | about 0.5 microseconds across the wafer | USD 50,000 per GB | by the system only | 0.6 | 0.1x to 0.4x: a uniformly random dependent chain is bisection-bound on a 2D mesh (about 110 G reads per second per wafer, up to 500 G with the dataset replicated twenty times), 5 to 22 M reads per second per watt against the 5090's 54 M; 1/1,000 per dollar | none needed; layer 2 removes the replicas as the dataset grows |
| CXL memory pools (CXL 2.0 and 3.x; about 70 ns of controller on top of local DDR5, 100 to 160 ns on Xeon 6 against 75 local, claimed, Introl, February 2026; USD 4 to 7 per GB before the 2026 cycle) | now | DDR5's 2 to 3 nJ plus the link | 170 to 250 ns through a host | USD 4 to 21 | anyone | 0.2 to 0.3 | 0.4x: 500 M 64-byte transactions per second per x8 device at about 25 W (approximate) | none needed; a capacity tool, not a random-read engine |
| Optical I/O: Ayar Labs TeraPHY (UCIe-compliant, about 5 pJ per bit, USD 500 M raised 3 March 2026, claimed); Celestial AI Photonic Fabric (6.2 pJ per bit, about 120 ns round trip, claimed) | 2027 or later | 6x to 8x the interposer's 0.8 pJ per bit per bit moved | plus 120 ns | n/a | n/a | worse than copper for this traffic | no edge: a pooling fabric; the dependent chain wants the memory closer, not farther | none needed |
| Chiplets and die-to-die: UCIe 0.25 to 0.5 pJ per bit (claimed, UCIe via SNIA); BoW 0.5 to 0.7; CoWoS-S USD 600 to 900 per H100-class package, a one-stack interposer about USD 200 (the record) | now | adds 0.27 nJ per 544-bit read that crosses a die boundary | plus 5 to 10 ns per hop | the package: USD 100 to 200 organic with UCIe-S, USD 200 to 900 for 2.5D | UCIe IP from several vendors; CoWoS capacity booked by AI through 2027 (approximate) | n/a | moves the SRAM chip's multi-die rows (below) and lets a 28 nm controller sit beside memory on an organic package | none needed |
| SRAM scaling at 2 nm: TSMC N2 38 Mb/mm^2 HD, +11 percent over N3E (claimed, IEEE Spectrum, 12 December 2024); wafers about USD 30,000, booked to 2028 (claimed, 2026) | now (N2 in volume from 2025) | 1.0 nJ for a 64-byte read from a 452 mm^2 array (approximate; 0.5 to 2.0: the global wire at about 1.3 pJ per bit across a 24 mm die is the term) | 10 to 20 ns | USD 200 to 300 per GB of silicon (one 600 mm^2 die, about USD 400 to 600, holds 2 GiB) | N2 is a merchant node: any customer with a project (Apple, AMD, NVIDIA, MediaTek, Qualcomm; the Bitcoin chip makers are on N3 and N4 class already, approximate) | 0.09 | **13x to 17x (4x to 5x)**: no activate ceiling, power-bound at about 2,100 MH/s per die at 300 W, USD 0.3 per MH/s of silicon; two dies at 4 GiB and four at 8 GiB cost USD 1,000 to 2,500 and read about 13x | **beats layers 1, 3 and 4; layer 2 cuts its capex, not its joules**: one reticle at 2 GiB, two at 4 GiB, four at 8 GiB; the shadow at k = 0.5 holds 4.8x, at k = 1 2.7x; on the M5 Max's joule about 3x at k = 0.5 |
Density against the schedule: SRAM gains 6 to 11 percent per node every two to three years (N3E +6, N2 +11, claimed),
the dataset doubles every four years (spec 1.13.3: 2 GiB at genesis, 4 GiB at year 4). The SRAM chip loses that race
slowly: its die count doubles each doubling and its per-read energy rises 0.1 to 0.3 nJ per extra hop. It does not lose
it inside five years.
## 3. What each line does to the four layers, read the other way
| Layer | What it does to the lines above | Number |
|---|---|---|
| 1. Per-era draws (mixer rounds, op-mix weights, read width, program length, shadow placement) | The read width and the mixer draws move no chip row: every chip in the table stores the dataset and reads the atom. The program length IS the shadow, and it is the only draw that reaches the memory-system chips | at `F` = 1.10 microjoules (class v4) the strongest DRAM chip reads 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5; at the full 5090 shadow (`F` about 2.0) 2.0x and 3.7x |
| 2. The state-sized dataset with a floor | Kills per-bank PIM and UPMEM outright; raises the SRAM chip's die count and capex; does nothing to any DRAM chip, which carries 24 to 64 GB per stack and over-provisions capacity for channels anyway | PIM serves 1.6 percent of reads at 2 GiB, 0.4 at 8; the SRAM chip's silicon USD 500 at 2 GiB, 1,000 at 4, 2,500 at 8, its joules 17x to 13x |
| 3. Scheduled family epochs every 180 days | None of the chips in the table is fixed-function; the controller chip runs any program and the SRAM chip's shadow core is a sequencer | 0 |
| 4. The (c''') floor and the F8 uniformity test per era | Keeps the hot-set cache at the record's 1.067x ceiling; the L2 hit at 2.4 nJ on the 5090 against 0.2 to 0.5 on a chip means the uniformity test is what stops a small cache from being the chip's edge | 1.067x at the ceiling (the record) |
| A fifth: the denominator | The honest tier's own operating point is the resistance: the 5090 at the lock reads 1.67 microjoules, the M5 Max 0.78; every chip edge halves to thirds against them | the SRAM die 5x and the base die 3.6x to 4.4x against the M5 Max at zero shadow; about 3x with the shadow at k = 0.5 |
## 4. The lines in detail
### 4.1 Processing-in-memory and processing-near-memory
Samsung's HBM-PIM (Aquabolt-XL, February 2021) places a programmable computing unit inside each bank of an HBM2 stack
and reports 2.5x system performance and over 60 percent energy saved on a Xilinx Alveo host (claimed, Samsung, 24
August 2021). SK hynix's GDDR6-AiM (February 2022) adds compute to a 16 Gbps GDDR6 die at 1.25 V, claims up to 16x on
some AI operations and 80 percent less power, and the AiMX card (2023) carries 32 GB of it (claimed, SK hynix). Samsung's
LPDDR5X-PIM (Hot Chips, 25 August 2026) reaches 614 GB/s inside the package against 76.8 GB/s over the external
interface at LPDDR5X-9600, 3x the tokens per second on Llama 3.1 8B, with mass production "as early as 2027" and
LPDDR6-PIM in JEDEC work (claimed, TrendForce and Sammyfans, 25 to 26 August 2026). UPMEM's DPU-in-DRAM ships: 128
DPUs per 8 GB DIMM, 64 MB per DPU, 350 MHz, MRAM reached only by DMA with a fixed cost plus a per-byte cost and at most
628 MB/s per DPU for 2,048-byte transfers, and "there is no direct communication channel among DPUs" (claimed, the PrIM
characterisation paper, arXiv 2105.03814, read 8 October 2026). The academic line (IMPICA, ICCD 2016: a pointer-chasing
engine on a 3D stack's logic layer, 1.2x to 1.9x and 10 to 41 percent less energy on linked lists, hash tables and
B-trees, claimed) puts the chaser on the base die, not in the bank, for the same reason this hash defeats the bank
units: the next address is anywhere.
What it does to the chip's cost per dependent read: nothing good for the attacker at the bank level. The PIM unit's
arithmetic is the wrong kind (FP16 SIMD, not 32-bit integer ARX) and the wrong place: with uniform addresses the chain
leaves the bank after one read with probability 1 minus bank bytes over dataset bytes (98.4 percent at 2 GiB on a 32 MB
bank), and in UPMEM's case leaves the DPU with no path but the host. The near-memory version (a chaser with its own
controller on the base die) is the record's `f = 1` chip, and that is where PNM becomes real: see 4.2.
Blunted by: layer 2, by the dataset's size alone. k: none (no path).
### 4.2 HBM3E, HBM4 and the custom base die
JEDEC's JESD270-4 (16 April 2025; read via eeNews Europe, 18 April 2025, and the search summaries, the Business Wire
and All About Circuits pages refusing the fetch) doubles the channel count to 32 with two pseudo-channels each on a
2,048-bit interface at up to 8 Gbps, 2 TB/s per stack, 4 to 16-high stacks of 24 or 32 Gbit dies up to 64 GB, VDDQ 0.7
to 0.9 V and VDDC 1.0 to 1.05 V (claimed). TSMC builds the standard HBM4 base die on N12 at 0.8 V for "roughly 1.5x"
efficiency and the custom C-HBM4E base die on N3P at 0.75 V for "2x the power efficiency of today's DRAM
manufacturing", and "the custom base die will integrate memory controllers and PHY components typically housed
separately"; Micron targets 2027 production, SK hynix a first tailored HBM4E in the second half of 2026 with 12 nm for
mainstream and 3 nm for NVIDIA and Google premium designs, Samsung 4 nm now and 2 nm for custom HBM (claimed,
TrendForce, 1 December 2025 and 23 January 2026). Prices: HBM2e USD 120 per 16 GB, HBM3 200 per 24, HBM3E 300 per 36,
HBM4 about 550 per 36 (estimate), factory gate, with contract about 2x and 20 to 26 week leads (claimed, Silicon
Analysts, October 2026); Samsung is asking "mid-to-high $4 per gigabit" for 2027 HBM4 against about 1.5 for HBM3E
(claimed, Sammyfans, 2 October 2026).
What it does to the chip's cost per read: the row activation (909 pJ per 1 KB row, HBM2, O'Connor Table 3) does not
move; the movement and I/O terms fall with the voltage and the base-die node; a 32-channel stack doubles the activate
parallelism the record's HBM rows are bound by (10.7 G reads per second per HBM3 stack, unmeasured; 2.3 G at the JEDEC
tFAW). The arithmetic on the record's method: HBM4 1.0 to 1.1 nJ per read and 21.3 G reads per second per stack (166
MH/s at about 36 W, 0.22 microjoules, 11x; at the JEDEC-tFAW ceiling 36 MH/s at 19 W, 0.53, 4.5x); the custom base die
0.9 to 1.0 nJ with its controller inside (166 MH/s at 29 W, 0.18, 14x; 6.5x at the low ceiling). The latency stays in
the 45 to 50 ns class; the capacity (36 to 64 GB) is 4x to 8x any floor layer 2 could set without retiring the honest
cards.
Who can buy it: through 2027 the stacks are allocated to AI accelerators at Tier-1 volume terms; a custom base die is a
per-customer engagement with the DRAM maker. A chip maker of Bitmain's or Canaan's size (the history's rows: tape-outs
on 7 nm-class nodes, R&D in the tens of millions of dollars a year) could commission one from 2028 (approximate); a
USD 5 M startup cannot. The brake is money and queue, not physics, and it expires.
Blunted by: nothing among the four. The shadow at `k` = 0.5 holds 4.4x, at `k` = 1 2.6x; the shadow core is logic on
the N3P base die and sits at the band's low end against a GPU on an older node.
### 4.3 LPDDR6
JESD209-6 (2025) gives 10,667 to 14,400 MT/s on a 24-bit channel split into two 12-bit sub-channels, a 32-byte minimum
access with 32 and 64-byte bursts, 4 to 64 Gb dies, lower voltages than LPDDR5, a dynamic efficiency mode and on-die
ECC (claimed, JEDEC via PCWorld, HotHardware and MicrocontrollerTips, 2025). The bank count per channel and the tFAW
are not in the public summaries read (unverified); LPDDR5's 16 banks per channel and a tFAW near 20 ns are the
assumption (approximate). A 32-channel, 64-sub-channel controller chip then reads about 12.8 G dependent reads per
second (approximate), 100 MH/s, at 1.5 to 2.0 nJ per read: about 4x to 5x the 5090 per joule, the GDDR7 board's class
at a lower price per channel. The dies are the cheapest random-access memory a non-hyperscaler can buy outside the
2026 price cycle (USD 3 to 4 per GB in 2024, approximate; USD 8 to 21 in 2026 LTAs, claimed).
The point of this line is not the chip. The M5 Max already is an LPDDR5X part with a GPU, at 0.78 microjoules per hash
measured at the GPU-plus-DRAM meter, 3.1x the 5090 per joule; the Windows-class LPDDR5X SoCs (NVIDIA's and Qualcomm's
desktop parts, approximate) and the LPDDR6 generation after them are the honest tier's floor. A controller chip on the
same memory beats that tier by 1.3x to 1.6x at zero shadow. The resistance of the chain is set by this tier, and v6
should say so (finding 3).
Blunted by: none needed.
### 4.4 3D DRAM
Samsung's VS-DRAM (VLSI 2023), 16 stacked layers demonstrated against Micron's 8, VCT 4F2 cells as the stepping stone
with prototypes in 2025 and commercial 3D DRAM "by approximately 2030"; Neo Semiconductor's 3D X-DRAM at 230 layers
and 128 Gbit per die as a concept (claimed, heise 29 May 2024, Yole, ComputerBase). No source read claims a latency or
activate-rate change; the gain is capacity per area (about 3x). For a chip that reads 1 to 8 GiB at random, capacity
per die is not the constraint (channels are), so 3D DRAM lowers the honest card's and the chip's cost per GB alike and
changes no row. Outside the five-year edge.
### 4.5 CXL memory pools
CXL 2.0 expanders on Xeon 6 measure 100 to 160 ns against 75 ns local DDR5; a CXL 3.1 controller adds about 70 ns;
pooled DDR5 was USD 4 to 7 per GB before the 2026 cycle (claimed, Introl, 1 February 2026, and the search summaries).
A dependent chain through a host CPU and a PCIe-class link at 64 bytes per transaction is bound by the link's
transaction rate (about 500 M per second per x8 device at about 25 W, approximate): 0.4x the 5090 per watt. A capacity
tool. No edge, no layer needed.
### 4.6 Wafer-scale
The WSE-3 holds 44 GB of SRAM at 21 PB/s across 900,000 cores on 46,225 mm^2 of N5; the CS-3 draws about 23 kW and
costs "maybe $2.5 million" (claimed, Cerebras and The Next Platform, 14 March 2024; the power from the search summaries).
A 2 GiB dataset spread over the wafer is read by a dependent chain whose next address is uniformly random across 215
mm of mesh: the traffic is all-to-all, the mesh is bisection-bound, and the average reply crosses about 140 mm of wire.
At about 0.1 pJ per bit per mm (approximate) a 512-bit reply costs about 7 nJ before routers, 3x the GDDR7 board's
2.0; at about 950 links across the bisection at 32 bits per cycle and about 1 GHz (approximate), the wafer completes
about 110 G reads per second, 500 G with the dataset replicated twenty times in regions, 5 to 22 M reads per second per
watt against the 5090's 54 M. Per joule 0.1x to 0.4x, per dollar one thousandth. The wafer is a streaming machine; this
hash is not streaming. No layer needed; layer 2 removes the replicas as the dataset grows.
### 4.7 Chiplets, interposers and optical I/O
UCIe gives 0.25 to 0.5 pJ per bit by package type (claimed, UCIe consortium via SNIA SDC 2022 and 2025 pages); BoW 0.5
to 0.7; CoWoS-S is USD 600 to 900 per H100-class package and a one-stack interposer about USD 200 (the record). A
read that crosses a die boundary pays about 0.27 nJ (544 bits at 0.5 pJ) and 5 to 10 ns, which is what makes the
multi-die SRAM chip of 4.9 cost 1.1 to 1.3 nJ per read instead of 1.0. Optical I/O (Ayar Labs' TeraPHY at about 5 pJ
per bit, UCIe-compliant, a USD 500 M round on 3 March 2026; Celestial AI's Photonic Fabric at 6.2 pJ per bit and about
120 ns round trip; claimed) is 6x to 8x the interposer's energy per bit and adds latency: it pools memory across
packages, which this traffic never wants. No row moves.
### 4.8 FPGA with HBM
The Alveo V80 (Versal XCV80, 32 GB HBM2e as two 16 GB stacks, 820 GB/s, 190 W, USD 9,495 MSRP, May 2024) and the U55C
(USD 4,747) are what a non-hyperscaler can buy today with HBM on it; Altera's Agilex 7 M-series carries the same two
HBM2e stacks at 820 GB/s (claimed, AMD, Wccftech, The Next Platform). The record's reading stands: an HBM2 stack under a
soft controller reaches the JEDEC tFAW ceiling, 2.3 to 2.4 G random reads per second (Shuhai, FCCM 2020, Figure 7),
so two stacks give about 4.8 G, about 37 MH/s at about 150 W (approximate), 0.6x the 5090 per joule at 5x its price. No
HBM3E FPGA was found announced in the pages read (unverified). No layer needed.
### 4.9 SRAM at 2 nm: the full store on one reticle
TSMC's N2 reads 38 Mb/mm^2 of high-density SRAM, 11 percent over N3E (claimed, IEEE Spectrum, 12 December 2024;
volume from 2025); wafers are about USD 30,000 and "booked to 2028" (claimed, tech-insider, 2026). The arithmetic: 2
GiB is 17,180 Mbit, 452 mm^2 of macro; with periphery, a controller, the lanes' registers and a shadow core, a 550 to
650 mm^2 die under the 858 mm^2 reticle; about 95 gross dies per wafer, 55 to 70 percent good with row and column
repair (approximate), USD 400 to 600 of silicon. No DRAM, no interposer, an organic package. Under class v5 the
dataset's items are leaves of the chain state refreshed per window; the chip rewrites 2 GiB per window at on-die
bandwidth, the same 32 ms every GPU pays (the record, chip-model-v3 5.10), and holds a node or shares one across a farm
as the record prices.
Energy per read: the macro's 64-byte read, about 0.1 nJ (approximate), plus the global wire. The record took 0.6 pJ per
bit of wire across a 128 mm^2 array; scaling with the side of the die gives about 1.3 pJ per bit across 600 mm^2, 0.68
nJ for 512 bits, so about 1.0 nJ per read with the controller, range 0.5 to 2.0 (the sensitivity of every SRAM figure
here). Latency 10 to 20 ns. No activate ceiling, no tFAW, no refresh: the rate is power-bound. At 300 W with 30 W of
static and controller power: 270 W over 128 nJ per hash is about 2,100 MH/s per die, 0.14 microjoules per hash, 17x
the 5090 at zero shadow (13x at 1.3 nJ, 8x at 2.0, 30x at 0.5); USD 0.25 per MH/s of silicon, about 0.4 with the
board. Two dies at 4 GiB: half the reads cross one UCIe hop, 1.14 nJ, 15x, USD 1,000. Four dies at 8 GiB: 1.3 nJ, 13x,
USD 2,000 to 2,500, a 4-die organic package. The honest cards hold 8 GiB fine (16 GB and up); a floor that pushed the
dataset past what a package can hold (16 to 32 GiB, eight to sixteen reticles at USD 5,000 to 10,000 and 1.5 to 2.0
nJ per read, still 7x to 9x) would retire every honest card under 32 GB first. Layer 2 therefore sets this chip's
capex and die count, not its joules, and the floor's number is a card-lifetime decision, not a chip decision.
The project: the history's IBS figures (5 nm USD 416 M to 542 M, 3 nm 590 M) price an SoC; an SRAM array with a
controller and a sequencer core is simpler and the startup figure ("$50M to $75M" for 7 nm, SemiAnalysis) is the
better guide, so USD 100 M to 500 M at N2 (approximate) and 18 to 24 months to a first chip (approximate). That is the
clock of Counter ASIC 3.0 item 4 and the daily-issuance threshold of the history (chips at USD 20 K to 50 K of daily
issuance for compute-bound hashes; 32 months for Ethash at the largest prize): the SRAM chip arrives when the prize
pays for an N2 project, and nothing in the hash moves that date.
Blunted by: layer 2 on capex only (USD 500 to 2,500 across the floor's range); the shadow at `k` on joules: with the
shadow core on the same N2 die against a GPU on N3 or N4 the ALU band sits at 0.3 to 0.5; at `F` = 1.10 the chip reads
4.8x at k = 0.5 and 2.7x at k = 1; at the 5090's whole latency shadow (about 330,000 ops per hash, `F` about 2.0,
575 W) 3.7x and 2.0x; against the M5 Max's joule about 3x at k = 0.5.
### 4.10 RLDRAM and the activate-free line
RLDRAM 3 (Micron, 2011; ISSI second-sourced) has a tRC under 10 ns, 16 banks per device, no separate activate
command and "SRAM-like random access" (claimed, Micron's product page and the 2011 announcements), in 576 Mb and 1.125
Gb devices at up to 2,133 Mb/s (approximate). One device completes about 2 G random reads per second (16 banks over 8
ns), six times a GDDR7 device's share of the 5090 board's 21.3 G; sixteen devices, 2 GiB, about 32 G reads per second,
250 MH/s per board. The energy per read is not published in anything read (a power calculator exists); a full
small-row access per read at an old node reads 2 to 4 nJ (approximate), so about 5x per joule, at USD 400 to 800 per
GB (approximate, a low-volume networking part). It is the proof that an activate-free DRAM exists, and AMD's "Folded
Banks" (ISCA 2025, with AMD Research: 8x the activate parallelism in a 3D-stacked HBM gives 6.7x the irregular
bandwidth, claimed from the abstract; the PDF refused the fetch) is the same idea on the HBM roadmap. If a DRAM maker
ships it in HBM, the activate ceilings in the record's HBM rows rise 6x to 8x and the HBM chip's rate per stack with
them; its energy per read falls by the row-size term (O'Connor's FGDRAM: 51 percent). That is the DRAM-on-logic row.
Blunted by: layer 2 on RLDRAM's capacity cost (USD 3,000 to 6,000 at 8 GiB); nothing on the HBM version.
## 5. The k bands for the research lane's chip rows
For `docs/design/class-v6-rotating-family.md`. The shadow `k` is the chip core's energy per forced op over the GPU's
at the same operating point (the record's ALU band 0.3 to 0.8 on the 5090's measured 6.2 to 11.3 pJ per op); the
memory `k_read` is the technology's energy per dependent read over the 5090's measured 10.9 nJ. "Edge" is per joule at
zero shadow; "with the shadow" is at the class v4 premium `F` = 1.10 microjoules on the 5090 (`E_chip` = `E_mem` + `k F`).
| Chip row | Year | `E_read` nJ | Rate per chip, MH/s | `E_hash` at zero shadow, microjoules | Edge over the 5090 (2.40) | Edge over the M5 Max (0.78) | `k_read` | Shadow `k` band | With the shadow, k = 0.5 / 1 | Silicon and memory, USD per MH/s |
|---|---|---|---|---|---|---|---|---|---|---|
| GDDR7 board, 28 nm controller (the record) | now | 2.0 | 166 | 0.466 | 5.1x | 1.7x | 0.18 | 0.5 to 0.8 (28 nm core: the high end) | 3.3x / 2.1x | 2.8 (2025 memory), about 7 in the 2026 cycle |
| HBM3E, one stack | now | 1.2 | 84 (ceiling unmeasured) | 0.321 | 7.5x | 2.4x | 0.11 | 0.3 to 0.8 | 3.8x / 2.4x | 6.6 |
| HBM3E, eight stacks | now | 1.2 | 666 | 0.262 | 9.2x | 3.0x | 0.11 | 0.3 to 0.8 | 4.1x / 2.5x | 4.0 |
| HBM4, one stack, N12 base die | 2027 to 2028 | 1.0 to 1.1 | 166 (36 at the JEDEC tFAW) | 0.22 (0.53) | 11x (4.5x) | 3.6x (1.5x) | 0.10 | 0.3 to 0.8 | 4.4x / 2.6x | about 5 |
| Custom HBM4E base die, N3P, controller in the stack | 2028 or later | 0.9 to 1.0 | 166 (36) | 0.18 (0.37) | 14x (6.5x) | 4.4x (2.1x) | 0.09 | 0.3 to 0.5 (an N3P core) | 4.4x / 2.6x | about 5 |
| DRAM on logic, FGDRAM-class rows, hybrid bonded | 2029 to 2031 | 0.5 to 0.7 | 125 per stack, 1,000 for eight | 0.15 | 16x (12x to 20x) | 5.2x | 0.05 | 0.3 to 0.5 | 5.1x / 2.8x | about 4 (approximate) |
| SRAM full store, one N2 reticle, 2 GiB | 2027 to 2028 (an N2 project) | 1.0 (0.5 to 2.0) | about 2,100 at 300 W | 0.14 | 17x (8x to 30x) | 5.6x | 0.09 | 0.3 to 0.5 (an N2 core) | 4.8x / 2.7x | 0.25 to 0.4 |
| SRAM full store, two N2 dies, 4 GiB | the same | 1.14 | about 1,850 | 0.16 | 15x | 4.9x | 0.10 | 0.3 to 0.5 | 4.6x / 2.7x | 0.55 |
| SRAM full store, four N2 dies, 8 GiB | the same | 1.3 | about 1,600 | 0.185 | 13x | 4.2x | 0.12 | 0.3 to 0.5 | 4.5x / 2.6x | 1.3 to 1.6 |
| LPDDR6 controller chip, 32 channels | 2027 | 1.5 to 2.0 | about 100 | 0.25 to 0.30 | 4x to 5x | 1.3x to 1.6x | 0.15 to 0.18 | 0.3 to 0.8 | 3.1x / 2.0x | about 3 (approximate) |
| Per-bank PIM, UPMEM, FPGA with HBM2e, wafer-scale, CXL, optical | | | | | under 1x or no path | | | | | |
Arithmetic, the SRAM row: 270 W over (128 x 1.0 nJ) = 2.11 G hashes per second; 300 W over 2.11 G = 0.142
microjoules; 2.40 over 0.142 = 16.9x; with the shadow at k = 0.5: (0.142 + 0.55) over (2.26 + 1.10) = 0.692 over 3.36 =
4.86x; at k = 1: 1.242 over 3.36 = 2.71x. The base-die row: 166 M x 128 x 0.95 nJ = 20.2 W plus 4 static plus 5
controller = 29.2 W; 29.2 over 166 M = 0.176 microjoules; 2.40 over 0.176 = 13.6x. The M5 Max column divides 0.78 by
the same `E_hash`. Every chip-side figure is modelled; the GPU-side figures are the record's measurements.
## 6. Consequences per tier
| Tier | What this file means | What is being done |
|---|---|---|
| Home miner, one 8 GB card | Nothing changes today: no chip exists, and the first one in this file (an N2 SRAM die or an HBM4 base-die chip) is a USD 100 M-class project with a 2028-class date. When one lands it runs at 0.14 to 0.22 microjoules per hash against this card's 10 to 20; this tier is the first out, as the record says. The dataset's floor decides this tier's life more than any chip does: 4 GiB at year 4 (the spec's schedule) retires it then | the shadow sizing and the floor are the founder's numbers to set (section 7); the share-pattern detector (Counter ASIC 3.0 item 4) is what tells this miner a chip has arrived |
| One 16 GB card (9070 XT class) | Holds 8 GiB with room; AMD's 2.4 G reads per second at 304 W is 7x behind the 5090 per joule and 50x to 100x behind the chips here | the vendor-share metric; nothing in the hash moves AMD's dependent-read rate |
| One 24 or 32 GB card, the 5090 at the lock | 1.67 microjoules at the 1,300 MHz lock; the chips here are 8x to 12x ahead per joule at zero shadow and 2.5x to 4x with the class v4 shadow at k = 0.5 | the Ember knob carries the lock rows; the shadow's upper bound at the full ALU budget is the lever this file sizes |
| The unified-memory SoC (M5 Max, LPDDR5X and LPDDR6 desktops) | 0.78 microjoules at the GPU-plus-DRAM meter: the honest tier the chips beat least (1.7x to 5.6x at zero shadow, about 3x with the shadow). Capex-poor (27 MH/s per USD 4,000 machine) but joule-rich | finding 3: v6 scores the floor and the shadow per tier with this tier as the reference joule; the dataset stays inside 16 GB unified memory |
| A rig | A rig's cost is electricity; against an N2 SRAM chip at USD 0.3 per MH/s and 0.14 microjoules it earns 1/10 to 1/17 of a chip per watt and leaves when chips hold the hashrate | the issuance trigger: the bounty and the benchmark live before daily issuance crosses about USD 50 K (the record) |
| A pool user | A chip fleet is a few operators; the share-pattern detector is the warning | the detector on the observer, before the public testnet (the record) |
| The public claim | "Under 2x" is not reachable against any chip in this file at a `k` under 1. The honest sentence is: the strongest chip five years out beats a 5090 by 2.6x to 2.8x per joule at k = 1 and about 4.5x at k = 0.5 with the shadow on, and the best honest SoC by about 3x; and it costs an N2 project | the research lane's synthesis carries the number; nothing from this file goes to the site or the devnet |
## 7. Decisions this raises for the founder
Each carries a default and a deadline; silence means the default.
1. **Size the shadow against the SRAM chip, not the GDDR7 board.** Layer 1's program-length draw gets a lower bound at
the length that holds the record's 2.1x on the GDDR7 chip today and an upper bound at the honest cards' full latency
shadow (the 5090's about 330,000 ops unlocked, about 150,000 at the lock; the M5 Max about 290,000; the 9070 XT
about 650,000; approximate from the record), re-based at every family epoch on the cards then mining. Default: the
research lane writes the draw with these bounds into the synthesis; the hash lane measures the rate and the watts at
the upper bound on the 5090 and the M5 Max before the first v6 era is cut. Deadline: 20:00 UK today (the synthesis).
2. **The floor (layer 2).** 2 GiB at genesis is one N2 reticle; 4 GiB two; 8 GiB four. The floor moves the chip's
capex (USD 500 to 2,500), not its joules (17x to 13x), and 8 GiB is the last size inside 16 GB unified memory and
the 16 GB card tier. Default: the spec's schedule stands (2 GiB plus 0.5 GiB a year, doubling at year 4), chosen on
card lifetime; the floor is not a chip lever and this file does not ask to raise it. Deadline: none; a note in the
synthesis.
3. **The reference joule (a fifth layer, or a rule).** Resistance is stated against the best honest joule (the
unified-memory SoC tier, then the 5090 at the lock), not the unlocked 5090. Default: the research lane adopts it in
the synthesis's chip rows (section 5's M5 Max column) and the public text, when there is one, carries the edge over
the best honest joule. Deadline: 20:00 UK today.
4. **The clock.** The chips in this file are USD 100 M-class projects with 2028-class dates; the issuance trigger, the
benchmark and the share-pattern detector of Counter ASIC 3.0 item 4 are what decide when they are built. Default:
unchanged from the record. Deadline: none.
## 8. Unverified and owed
- Every chip-side energy figure is modelled on the record's method (O'Connor's HBM2 breakdown, Samsung's pJ per bit
roadmap, the 909 pJ row activation standing in for HBM3, HBM4 and GDDR7); the SRAM wire figure (0.6 pJ per bit at 128
mm^2, scaled with the die's side) is from memory and moves the SRAM rows by 2x either way; the UCIe hop (0.5 pJ per
bit) is the consortium's claim; the FGDRAM row size (256 bytes) and its 51 percent are the paper's simulation.
- HBM4's banks per pseudo-channel and tFAW per channel are behind the JEDEC paywall (the Business Wire and All About
Circuits pages refused the fetch); the "2x the activate ceiling" rests on 32 channels with tFAW per channel, as the
record's HBM rows rest on the same assumption at 16; both unmeasured. The AWS F2 hour the record names (chip-model-v3
5.3) is still the one measurement that would settle the HBM2 figure, and an HBM3 or HBM4 part is not rentable at the
controller level by anyone outside a hyperscaler today.
- LPDDR6's bank count and tFAW are not in the summaries read; the row is LPDDR5's structure (approximate).
- RLDRAM 3's energy per read and its 2026 price are not published in anything read; the row is approximate.
- The Cerebras mesh figures (link width, clock, bisection) are approximate; the fabric bandwidth was not on the pages
read (The Next Platform gives only its change against WSE-2).
- Dates: the DRAM-on-logic and hybrid-bonded HBM timelines ("2028 to 2031") are approximate; the sources read give
3D DRAM "from 2030" and C-HBM4E production in 2027; the hybrid-bonding pages were not fetched (the TrendForce page
refused). The N2 project cost and time are approximate.
- The web-search budget of this session ran out at about 11:3x UK after 26 searches; the remaining facts were read by
direct fetch of the pages named in section 9. Pages that refused (403): JEDEC's Business Wire release, All About
Circuits, ACM's Folded Banks page, OC3D's GDDR7 capacity note, All About Circuits' HBM-PIM note. Their figures are
carried from the search summaries and marked claimed.
- Nothing was run on the Mac; nothing was built or benchmarked anywhere. The one measurement this file would want next
is on the record's queue already: the 5090 and the M5 Max at the shadow's upper bound (decision 1).
## 9. Sources (URL and the date read; all read 8 October 2026 unless a file is named)
Repository: `docs/analysis/chip-model-v3.md` (sections 5.1 to 5.11, 6); `docs/analysis/counter-asic-4-research.md`
on branch counter-asic-4 at fb61ed4b (sections 1, 15.1a, 20.4); `docs/analysis/asic-resistance-history.md` (2.5, 2.6);
`docs/plans/counter-asic-3-status.md` (7c, the clock grid); `docs/spec/01-lottery-hash.md` (1.8, 1.13.3, 1.14);
`docs/analysis/latency-shadow-2026-10-06.md`; `docs/analysis/sram-mirror.md`.
- JEDEC HBM4 (JESD270-4, 16 April 2025): https://www.eenewseurope.com/en/hbm4-standard-doubles-channel-count-for-ai-boost (18 April 2025); https://hothardware.com/news/jedec-finalizes-hbm4-spec; https://www.businesswire.com/news/home/20250416843598/en (refused)
- TSMC HBM4 and C-HBM4E base dies: https://www.trendforce.com/news/2025/12/01/news-tsmc-unveils-custom-c-hbm4e-details-n3p-logic-dies-reportedly-target-2x-efficiency-gain/ (1 December 2025); https://www.trendforce.com/news/2026/01/23/news-samsungs-custom-hbm4e-design-reportedly-aimed-for-mid-2026-parallels-sk-hynix-and-micron/ (23 January 2026)
- HBM prices: https://siliconanalysts.com/data/hbm-pricing (October 2026); https://www.sammyfans.com/2026/10/02/samsung-seeks-more-than-3x-hbm3e-pricing-for-hbm4/ (2 October 2026); https://www.trendforce.com/news/?p=62555
- GDDR7 prices and roadmap: https://www.trendforce.com/news/2026/09/24/news-micron-reportedly-ends-2gb-gddr7-narrowing-supply-options-for-nvidias-rtx-50-series/ (24 September 2026); https://www.guru3d.com/story/micron-fiveyear-roadmap-shows-24gb-36gbps-gddr7-in-2026; https://overclock3d.net/?p=314548 (refused; the 4 and 6 GB parts in 2027 to 2028 from the search summary)
- DRAM prices 2026: https://wccftech.com/mobile-dram-prices-expected-to-increase-by-100-quarter-over-quarter-as-long-term-agreements-now-getting-signed-at-prices-as-high-as-21-gb/ (4 May 2026); https://tech-insider.org/ddr5-ram-prices-2026/
- Samsung HBM-PIM: https://news.samsung.com/global/samsung-brings-in-memory-processing-power-to-wider-range-of-applications (24 August 2021); https://www.allaboutcircuits.com/news/beyond-high-bandwidth-memory-samsung-breaks-processing-in-memory-into-AI-applications/ (refused)
- SK hynix AiM and AiMX: https://news.skhynix.com/developed-processing-in-memory (2022); https://www.hc2024.hotchips.org/assets/program/conference/day1/11_HC2024.SKhynix.GuhyunKim.rev920240822.pdf
- Samsung LPDDR5X-PIM and LPDDR6-PIM: https://www.sammyfans.com/2026/08/25/samsung-unveils-lpddr5x-pim-dram/ (25 August 2026); https://www.trendforce.com/news/2026/08/26/news-samsungs-4nm-gaia-could-mark-first-pim-commercialization-in-ai-pcs-mass-production-as-early-as-2027/ (26 August 2026); https://en.fnnews.com/news/202609060934024282
- UPMEM: https://arxiv.org/pdf/2105.03814 (the PrIM characterisation, text extracted with pdftotext: Table 1, section 3.2, the inter-DPU statement); https://www.nextplatform.com/2020/02/04/putting-in-memory-processing-through-the-paces/ (4 February 2020); https://old.hotchips.org/hc31/HC31_1.4_UPMEM.FabriceDevaux.v2_1.pdf
- IMPICA: https://ghose.cs.illinois.edu/papers/16iccd_impica.pdf (ICCD 2016)
- Fine-grained DRAM: https://www.cs.utexas.edu/~skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf (text extracted with pdftotext: 3.92 pJ per bit per HBM2 access, 1.21 of it activation, the 256-byte row, tFAW "effectively eliminated", 51 percent)
- Folded Banks (ISCA 2025): https://dl.acm.org/doi/10.1145/3695053.3731111 (refused today; the abstract as the record read it on 6 October 2026); https://wantongli.ucr.edu/news/announcement22-ISCA%202025
- LPDDR6: https://www.pcworld.com/article/2845759/lpddr6-memory-standard-announced-as-ddr5-dram-takes-over.html; https://hothardware.com/news/jedec-lpddr6-standard-released; https://www.microcontrollertips.com/what-is-jesd209-6-and-why-is-it-important-for-edge-ai/
- CXL: https://introl.com/blog/cxl-memory-expansion-pooling-disaggregated-memory-ai-data-center-2025 (1 February 2026); https://www.snia.org/sites/default/files/2025-09/SNIA-SDC25-Peethambaran-CXL-as-scalable-cost-effective-Memory.pdf
- Cerebras: https://www.cerebras.ai/chip; https://www.nextplatform.com/2024/03/14/cerebras-goes-hyperscale-with-third-gen-waferscale-supercomputers/ (14 March 2024); https://sacra.com/c/cerebras-systems ("a couple million per system")
- FPGA with HBM: https://wccftech.com/amd-announces-mass-production-of-the-alveo-v80-compute-accelerator-9495-price-tag/ (17 May 2024); https://www.amd.com/en/products/accelerators/alveo/u55c/a-u55c-p00g-pq-g.html; https://www.nextplatform.com/2022/03/08/a-cornucopia-of-memory-and-bandwidth-in-the-agilex-m-fpga; https://arxiv.org/abs/2005.04324 (Shuhai)
- 3D DRAM: https://heise.de/en/news/Huge-RAM-3D-DRAM-with-multiple-layers-planned-from-2030-9738064.html (29 May 2024); https://www.yolegroup.com/industry-news/samsung-reveals-16-layer-3d-dram-plans-with-vct-dram-as-a-stepping-stone/
- RLDRAM 3: https://www.micron.com/products/memory/dram-components/rldram-memory; https://newelectronics.co.uk/content/news/micron-unveils-third-generation-rldram-technology (2011); https://arxiv.org/pdf/1810.07059
- SRAM at N2: https://spectrum.ieee.org/tsmc-n2-2670436570 (12 December 2024); https://marklapedus.substack.com/p/intel-tsmc-tout-sram-breakthroughs; wafer prices https://tech-insider.org/tsmc-2nm-wafer-price-2026/
- UCIe and BoW: https://www.snia.org/sites/default/files/2025-05/SNIA-SDC22-Sharma-Universal-Chiplet-Interconnect-Express.pdf; https://www.3dtested.com/news/new-ucie-chiplet-standard-supported-by-intel-amd-and-arm
- Optical: https://www.theregister.com/2026/03/03/ayar_labs_500m/ (3 March 2026); https://www.allpcb.com/allelectrohub/photonic-interconnects-aim-to-solve-ai-memory-bottlenecks (the Celestial AI figures)
- H100 pointer-chase latency (353 ns, claimed): https://arxiv.org/pdf/2608.15764

View file

@ -0,0 +1,333 @@
# Class v6 research lane A, history: every ASIC-resistant proof of work, how it fell or held, and what that binds on the four layers
8 October 2026, branch `class-v6-history`, the Counter ASIC coordinator's history lane. The founder's word at 11:1x UK: class v6 is declared with four layers as its spine, and the research opens "see if anything can be optimised, added or invented". This file extends and corrects `docs/analysis/asic-resistance-history.md` (5 October 2026, the deep dive: 31 rows, 24 papers, ten lessons, seven ranked additions); it does not repeat that file's rows. What is new here: the exact MECHANISM of every chip (what it specialised: the memory, the hash core, the instruction mix, the parameter fixity), what each design missed, the timeline from announcement to chip to response, and for each one the mapping to class v6's four layers with "does v6 close it" in one sentence and one number. Every figure about another chain cites a URL with the date it was read, or is labelled approximate. Every Igneum figure names the repo file. Reading public research is in-house; nothing is paid or asked of anyone outside.
First cut landed on the mirror's master at 11:26 UK on 8 October (merge 4c58ad65: the table, the mechanisms, the three lessons), ahead of the 15:00 clock; this revision folds in the synthesis lane's reading of the six items (`docs/design/class-v6-rotating-family.md` section 7b) and cross-cites lane B's hardware file; the full report by 09:00 UK on 9 October carries any row the other lanes move. A line goes to the coordinator and to the synthesis lane at each landing.
## 0. One page
**The four layers, as declared, against the history's chip classes.** Every chip that ever beat a resistant hash belongs to one of five classes; the table says which layer of v6 answers each class, and with what number. "Closes" means the chip's measured or claimed edge falls under 2x per joule by a mechanism the layer supplies; "does not close" means the layer does not touch the chip's edge, and the number stands.
| Chip class (the rows of section 1 it covers) | What the chip specialised | Its best measured edge | Which v6 layer answers it | Does v6 close it | The number after v6 |
|---|---|---|---|---|---|
| A. Fixed-function pipeline on a compute-bound hash (X11, Blake-256, Blake2b, Blake2s, kHeavyHash, Eaglesong, Blake3, SHA512/256d, NexaPow, X16R's FPGA) | the hash's round function unrolled in silicon; the instruction mix fixed at genesis | 33x to 1,280x per joule (Antminer KA3, D3, KS5 Pro) | layers 1 and 3 (the op mix, program length and families drawn or unlocked per era; already closed since class v2 by the per-epoch random program) | yes, and it was already closed: a per-epoch program has no round function to unroll | 0 of 128 loads and 0 of 512 program ops are a fixed function; what remains is class E |
| B. SRAM-scale memory (Scrypt's 128 KB, CryptoNight's 2 MB, Lyra2REv2's sponge, Cuckatoo's edge bitmap, Equihash's 144 MB) | the hash's whole working set on die or across a few dies, with a time-memory trade-off the designer had not drawn | 19x to 1,100x (Scrypt), 40x to 50x (CryptoNight X3), 12x to 100x (Equihash Z9 to Z15), 4x (Cuckatoo G1) | layer 2 (the dataset above any SRAM die, growing with chain state with a floor; the 256 MiB cache growing with it) | yes: the recompute chip that holds the cache on die reads 0.92x at the op budget and 1.86x per joule, and the cache's size forces it onto a 7 nm or better node | `chip-model-v3.md` section 2 and 5.4: 0.31x bare, 0.92x with the 3x factor, 1.86x per joule at f = 0; the curve from f = 0 to 1 is monotone and the partial-store chip is worse than both ends |
| C. The memory system without the GPU (the Ethash chips: Linzhi Phoenix, Jasminer X4, Antminer E9; the f = 1 chip of the model) | a controller and PHY for the same DRAM the card carries, a 32-byte atom per read, no shader, no scheduler, no clock tree; 55 W of memory without the 271 W of GPU | 2.1x to 4.8x per joule measured (Ethash, 2020 to 2022); 5.1x (GDDR7) to 7.5x (one HBM3 stack) modelled for Igneum at zero premium | none of the four directly: every per-era draw is firmware to a chip that stores the dataset; layer 2 only when the dataset passes the chip's board | **NO.** The f = 1 chip keeps 5.1x per joule on GDDR7 at zero premium and 2.1x with the class v4 shadow at k = 1; v6's draws move neither figure. What moves it is the honest card's own joules (the operating point, 3.6x at the 1,300 MHz knee) and the shadow's premium | 5.1x / 3.6x / 2.1x (zero premium unlocked / knee / knee with the shadow at k = 1), `counter-asic-4-research.md` sections 0 and 20.4; the chip's USD 2.8 per MH/s against the card's 14.7 |
| D. The firmware-survivable chip against periodic change (Monero's chips across four forks, Vorick's "survives forks at under a 5x hit", the X16R FPGA, the Equihash chip "able to follow parameter forks") | a programmable sequencer over the hash's op set; the per-block or per-fork change read as a configuration | chips back at 85 percent of Monero's hashrate four months after the v8 fork; 1.3x on X16R | layers 1 and 3 are the automatic form of the change those chains made by hand; they cost no governance event | partly: v6 closes the GOVERNANCE failure (no fork, no reset of the hashrate to a rentable size) and taxes die area for the reserve; it does not close the chip, because a reserve and a draw band readable at genesis are built in from day one | the draw's cost to the chip is a recompile per epoch and a new dataset mapping per era; the number is the "GPU without graphics" of ledger M1, which is class C's chip with a sequencer beside it: 5.1x less the shadow's k |
| E. The hot set and the steered address (Kik's 64-bit seed on ProgPoW, Dinur-Nadler on MTP, AP-F8-1 on class v4, the weak-day MUL draw) | a small SRAM serving the reads a biased draw or a cooperating node concentrates | 1.067x at v4's ceiling (0.52 to 4.6 percent of reads on 0.1 percent of items); a memory skip on ProgPoW 0.9.3; MTP from 2 GB to under 1 MB | layer 4 (the (c''') floor and the F8-form uniformity test generalised to every era's draw, with a redraw on failure) | yes, for the class the test models: the floor at 0.995 refuses every live hot set the in-house pass found (nine, 0.9809 to 0.9919) at 0.7 percent of candidates | `docs/spec/01-lottery-hash.md` 1.4.7.2 (class v5); the open residue is the shadow-written concentrations at 0.9992 to 0.9997, worth about 1.0004x to a chip |
**The three lessons that bind on v6** (section 3 has the evidence):
1. **The chip that stores the dataset is firmware-immune to every draw; only joules and memory growth move it.** Every per-era parameter (mixer rounds, op weights, read width, program length, shadow placement) and every family epoch is read by the f = 1 chip's sequencer as a configuration, as Monero's chips read four forks and X16R's FPGA read the per-block order. The number v6 inherits is 5.1x per joule at zero premium (2.1x with the shadow at k = 1), and the lever is the honest card's operating point and the shadow's premium, not the draw.
2. **Automatic change beats the human fork only where it costs the chip a redesign, and the one parameter that does is the memory.** Grin's six-monthly tweaks held because each was a new algorithm the lane was scheduled to retire; Monero's forks lost on the second lap; Ethash's DAG growth is the one scheduled change in the record that killed a shipped chip (the E3, when the DAG passed its 4 GB of DDR3). Layer 2 is that lesson made a rule, and its rate decides everything: at 2 GiB plus 0.5 GiB a year a 32 GB chip board outlives the chain, so layer 2 as declared ages out the honest 8 GB card before any chip unless the floor is set against DRAM cost per gigabyte, not against chain state.
3. **A steered or biased address pattern is always found after launch unless the test lives in the acceptance rule, and every new draw needs its own null.** ProgPoW's seed, MTP's blocks and class v4's lossy sources were the same attack three times; layer 4 puts the test where it must be, and its cost is one census per era draw on the node (2.2 s per candidate at 2^20 today) with the null re-derived for every drawn parameter, because the window model that defines "uniform" changes with the read width and the program length.
**What the history says to add** (section 5): a layer 2 floor and a per-tier ceiling stated before its rate (the synthesis lane has since made the ceiling layer 2's first constant); the issuance clock and the share-pattern detector (unchanged from the 5 October ranking, still unbuilt); the read width's draw band stopped where the honest card stops being latency-bound ({4 B, 16 B} on the measured warrant, nothing wider); the op-mix band bounded by the per-vendor energy table; the reserve ordered by what a sequencer chip cannot fold into firmware.
## 1. The chips, one row per mechanism
Columns: the hash; the chip (vendor, model); the mechanism (what it specialised); what the design missed; the timeline (hash live, to the chip, to the response); the v6 layer that answers it; does v6 close it; the number (the chip's measured or claimed per-joule edge, and what the layer leaves). Every ratio is arithmetic on the cited rate and watt figures of the chip and the best consumer GPU of its year, approximate by construction; section 2 carries the sources. "Closed at v2" means the per-epoch random program already removed the mechanism and v6 inherits it.
| Hash (chain) | Chip | Mechanism: what it specialised | What the design missed | Timeline | v6 layer | Closed by v6 | The number |
|---|---|---|---|---|---|---|---|
| Ethash (Ethereum) | Bitmain Antminer E3 | 18 chips with 4 GB of commodity DDR3; "a memory interface connected to a small compute engine" | the hash needs only a memory interface; and the chip's memory was under-sized | live Jul 2015; chip announced Apr 2018, shipped Jul 2018; died by DAG growth Mar to Oct 2020 | 2 (the only layer that reached it) | class C: no | 1.0x to 1.1x per joule; the DAG passed its 4 GB at 20 to 27 months |
| Ethash | Innosilicon A10, A10 Pro, A11 Pro | the same with faster, larger DRAM (6 to 8 GB) | as above | Jul 2018 to Dec 2021 | none | no | 1.4x to 2.5x |
| Ethash | Linzhi Phoenix (E1400) | 64 compute units and 72 mixers per board beside 4.4 GB of "decentralized memory inside its ASIC"; sized to the DAG | memory energy per bit is the whole cost; many small memories by many small cores cut it | announced Sep 2018; tapeout Sep 2019; samples Dec 2020; no mass production | none | no | 2.1x to 2.2x |
| Ethash | Jasminer X4 (Sunlune) | DRAM dies hybrid-bonded (wafer-to-wafer DBI) onto a 40/45 nm logic die, 5 GB per unit, 40 chips per server | off-package DRAM at 3 to 4 pJ per bit was the cap; bonding takes it under 1 | Jun 2021 chip; Oct 2021 ship; the Merge 11 months later | none | no | 5.1x (about 7.3 pJ per bit all in) |
| Ethash | Bitmain E9, E9 Pro | conventional, 6 to 7 GB | as the E3, four years of DRAM later | Jul 2022 (8 weeks before the Merge); Feb 2023 | none | no | 3.0x, 4.1x |
| ProgPoW (never on Ethereum; KawPow, FiroPoW, ProgPowZ, Quai) | none; Linzhi claimed 3x to 8x with no derivation; one academic VU35P FPGA at 5.6 MH/s | a sequencer over eleven ops reads the per-period program as configuration (unbuilt); the on-die cache chip priced at "<< 0.1x" energy (unbuilt) | the light-evaluation chip was left as a suggestion; a 64-bit seed (Kik) | EIP May 2018; audits Sep 2019; exploit Mar 2020; dead Mar 2020; no chip on any adopter in 6 years | 1, 3 (the period made an era draw), 4 (the seed class) | class D and E yes; the Rao chip is class C, no | claimed 1.1x to 1.2x for the compute chip; the memory-system chip never priced; Igneum's reads 5.1x |
| RandomX (Monero) | Bitmain Antminer X5 | many RISC-V chips on a board over commodity DRAM: an in-order machine for a fixed VM spec; insides never published | CFROUND's cost on x86, a scratchpad-only AES circuit, the latency slack of a fixed 256-instruction program (v2's own list) | live Nov 2019; private mining from about 2021; X5 Sep 2023; X9 announced Dec 2025 and withdrawn May 2026 (zero shipped); v2 released Mar 2026, activation pending | 1 (program length and op mix drawn, which v2 had to fork to change) | class D: yes for the class; the gap itself has no GPU analogue | 1.46x (X5, measured); 3.8x (X9, claimed, never shipped); 5.4x (Pinecone R1X, claimed, undelivered) |
| CryptoNight (Monero) | Bitmain Antminer X3 (180 x BM1700); Baikal Giant N; secret chips from early 2017 | one or two 2 MB scratchpads in on-die SRAM with a hardware AES round per chip (inference from 1.2 kH/s at 2.6 W per chip) | the working set was a 2014 L3, which is a die | live Apr 2014; secret chips about 33 months; X3 announced at 47 and bricked by the v7 fork before delivery; chips back inside 4 months of v8; none in CN-R's 9 months | 2 (the dataset above any die); 1 and 3 (CN-R's per-block program is Igneum's per-epoch one) | yes | 40x (X3 against a Vega 64); Igneum's f = 0 chip 1.86x per joule |
| Cuckatoo31 (Grin) | Obelisk GRN1 (cancelled) | one TSMC 16 nm die with 512 MiB of SRAM holding the node bits, lean mining; 150 GPS at 800 W per chip | the resource was SRAM size at 320 MB, and a die was designed for it | hash live Jan 2019; announced Jan 2019; sold Apr; cancelled Jul 2019 | 2 | yes (closed at v3) | 50x planned; the 2025 bounty: memory trades for time, not energy |
| Cuckatoo32 (Grin) | iPollo G1 (30 x 12 nm chips) | node bits on die, the 512 MB edge bitmap in serial DRAM (the design's geometry; the G1's own split unpublished) | half the working set stayed in DRAM | Dec 2020, 23 months | 2 | yes | 3x |
| Equihash 200,9 (Zcash) | Bitmain Z9 mini, Z9 (BM1740); A9; Z11; Z15 Pro | the whole 144 MB working set on one die ("about 128 MB", eDRAM or SRAM); the Fudan NDSS 2019 design: an on-die linear sorter with the lists in off-chip DDR4, parameter-independent | the 1,000x halving penalty had no proof; 144 MB fit a die; a chip was designed to follow any (n, k) | live Oct 2016; secret chips before May 2018; Z9 mini at 18 months; Z15 Pro at 80 | 2 (a 2 GiB set is not a die); the parameter-following chip is layer 1's warning | yes by layer 2 | 12x (Z9 mini) to 110x (Z15 Pro); the Fudan design 13x in simulation |
| Scrypt (Litecoin) | Innosilicon A2; KnC Titan; Antminer L3+ (BM1485); L7 (BM1489) | one 128 KB scratchpad of on-die SRAM per core, 12 cores per 28 nm chip (L3+), 7 nm with 480 chips (L7); Percival's own lookup gap (half the memory for 25 percent more work) | the scratchpad was a 2011 cache; the time-memory trade favoured the chip at 2x to 4x by the designer's note | live Oct 2011; A2 Apr 2014 (30 months); L3+ Jun 2017; L7 Nov 2021 | 2 | yes | 70x (A2) to 1,300x (L7); Igneum's f = 0 chip 1.86x |
| Scrypt-N (Vertcoin 2014) | none; dropped pre-emptively | an automatic N doubling a chip follows with more SRAM or a wider lookup gap | public and slow growth sizes the chip for years | Jan 2014 to Dec 2014 | 2 (the precedent) | n/a | no chip shipped |
| Lyra2REv2 (Vertcoin) | FPGA bitstreams (2018), then Dayun Zig Z1 | a sponge of about 200 KB per core on die (approximate); 28 nm | SRAM-scale memory inside a hash chain | live Aug 2015; FPGA 2018; Z1 Sep 2018 (37 months); fork Feb 2019 | 2 | yes | 20x |
| X16R (Ravencoin) | CVP-13 FPGA bitstreams; OW1 and SKC Turing R1 | sixteen fixed cores with per-block routing | the draw changed the order, not the hardware | live Jan 2018; FPGA named Sep 2018; 45 percent of blocks by Jul 2019; fork Oct 2019; chips "evident" again Jan 2020 | 1 (the warning) | the fixed-function lane closed at v2; the draw is governance | 0.75x (OW1) to 5x (R1) |
| MTP, Argon2d (Zcoin) | none; Dinur-Nadler's attack before launch | the prover steers the data-dependent addresses into under 1 MB at a 170x penalty | the attacker controlled the memory's contents | 2017 attack; live Dec 2018; replaced Oct 2021 | 4 | yes (the day key is a VDF of chain state; the floor refuses hot sets) | 2 GB to under 1 MB; Igneum's residue about 1.0004x |
| Yescrypt, yespower | none | L2-latency-bound sequential work; small prizes | n/a | 2014 on | 2 | n/a | no chip |
| KawPow, Autolykos v2, Octopus, Verthash, FishHash | none | DRAM-scale random reads on small prizes | untested at a prize that pays for a controller project | 2020 to 2024 | 2 | class C: not tested | no chip; the Ethash record is their test |
| kHeavyHash (Kaspa) | IceRiver KS0 to KS5L; Antminer KS3, KS5 Pro | a fixed 64 x 64 nibble matrix pipeline | compute with a matrix in it is the cheapest silicon there is | live Nov 2021; KS0 Sep 2023 (22 months); GPU share gone by late 2023 | closed at v2; the mm8 reserve's warning | yes | 170x to 720x |
| Eaglesong (Nervos) | Toddminer C1; Antminer K5, K7 | a fixed sponge pipeline | compute | 4 months | closed at v2 | yes | about 100x |
| Blake3 (Alephium); Blake2s (Kadena); Blake-256 (Decred); Blake2b (Sia); X11 (Dash); SHA512/256d (Radiant) | the 5 October file's rows | fixed pipelines | compute | 4 to 31 months | closed at v2 | yes | 33x to 6,500x |
| NexaPow (Nexa) | DragonBall A21 | a secp256k1 Schnorr signature per nonce in hardware | a wide multiplier chain is already what the GPU does well | live 2023; chip Dec 2024 | closed at v2 | yes | 2.2x: the one compute row where the chip's edge stayed small |
What the table says, sorted by class: class A (fixed pipelines) 33x to 6,500x, closed at v2; class B (SRAM-scale sets) 12x to 1,300x, closed at v3 by the dataset and the drawn curve; class C (the memory system) 1.0x to 5.1x measured on Ethash, modelled 5.1x to 9.2x for Igneum, NOT closed by any layer; class D (firmware-survivable sequencers) 0.75x to 5x on the program side, closed as governance and open as a chip; class E (steered addresses) closed by layer 4 for the class it models.
## 2. The chips, in depth
One section per proof of work. Each carries the mechanism, the miss, the timeline, the layer mapping and the number. Where a lane's return says "not found" the row says so.
### 2.1 Ethash (Ethereum, July 2015): the E3, the Innosilicon line, Linzhi, Jasminer and the E9
**The mechanism of the hash.** A DAG of 1 GB growing 8 MB per epoch of 30,000 blocks (about 0.7 GB a year), 64 random 128-byte reads per hash mixed by FNV, keccak at each end; bandwidth-bound by design. EIP-1057 states the miss in one sentence: Ethash "requires external memory due to the large size of the DAG. However that is all that it requires - there is minimal compute ... a custom ASIC could remove most of the complexity, and power, of a GPU and be just a memory interface connected to a small compute engine" (https://eips.ethereum.org/EIPS/eip-1057, read 8 October 2026).
**The chips, and what each specialised.**
| Chip | Announced, shipped | The mechanism, as far as any source states it | MH/s, W, MH per joule | Against the best GPU of its year (derived, approximate) | Source (read 8 October 2026) |
|---|---|---|---|---|---|
| Bitmain Antminer E3 | leaked March 2018, announced 3 to 4 April 2018 at USD 800 (five per customer), shipped 16 to 31 July 2018; later batches USD 1,800 | 18 Ethash chips on three boards with 4 GB of commodity DDR3; chip, node, controller and read width never published; Bitmain's own support called it "a 4G video card" whose DDR "is up to the upper limit"; Rao's audit classes it as the conventional strategy (compute in silicon, memory off chip) | 180 claimed, 190 to 200 shipped, at 800 W: 0.24 MH/J | 1.0x to 1.1x against a tuned GTX 1080 Ti (45 MH/s at about 200 W, 0.225); under Rao's best overclocked 2019 GPU (0.40) | https://cryptoslate.com/bitmain-e3-asic-ethereum-miner/ ; https://2miners.com/blog/asic-miners-for-ethereum-antminer-e3-vs-innosilicon-a10-eth-master-comparison/ ; https://coingeek.com/memory-limitations-prompt-bitmain-antminer-e3-to-halt-etc-support/ ; https://www.kryptex.com/en/hardware/nvidia-gtx-1080-ti/reviews |
| Innosilicon A10 ETHMaster, A10 Pro (6 GB), A10 Pro+ (7 GB), A11 Pro (8 GB) | A10 announced 23 July 2018 at 365, 432 and 485 MH/s (USD 3,800 to 5,000); A10 Pro June 2020; A10 Pro+ January 2021; A11 Pro presold March 2021 at 2,000 MH/s and 2,500 W, shipped December 2021 at 1,500 MH/s and 2,350 W ("20 percent less efficient"; broker quotes about USD 27,000) | the same conventional chip with a larger and faster memory system; the DRAM type is not stated on any page read (GDDR6 is the trade's assumption, approximate); node and read width not found | A10 485 at 850 W: 0.57; A10 Pro 500 at 860: 0.58; A10 Pro+ 750 at 1,350: 0.56; A11 Pro 1,500 at 2,350: 0.64 (0.80 claimed) | A10 2.5x against the 1080 Ti; A10 Pro 1.4x and A11 Pro 1.6x (2.0x claimed) against a tuned RTX 3090 (120 MH/s at about 295 W, 0.41) | https://www.criptonoticias.com/mineria/nuevo-minero-asic-innosilicon-procesa-485-mh-ethereum ; https://www.theblock.co/post/125871/innosilicon-ethereum-miner-a11-pos ; https://whattomine.com/coins/151-eth-ethash/asics |
| Linzhi Phoenix (E1400) | announced 13 to 14 September 2018 by Chen Min (ex-Canaan) at 1,400 MH/s and 1 kW for April 2019; planned tapeout December 2018, actual September 2019; sample rollout 21 December 2020; no mass production; price never published; the company now "studying new opportunities" | two E1400 boards, each 64 compute units and 72 "mixers" with 4.4 GB of "decentralized memory inside its ASIC design" (The Block); the DRAM type never published (one 2019 forum post says an interposer and stacked HBM dies, unverified); the 4.4 GB sized to the DAG rather than to a commodity 6 or 8 GB, which reads as memory sized per chip (inference); Linzhi claimed a ProgPoW chip would reach 3x to 8x | 2,600 claimed, 2,733 measured by F2Pool, at about 3,000 W: 0.87 to 0.91 | 2.1x to 2.2x against the tuned 3090 | https://www.theblock.co/post/88622/questions-new-ethash-asic-ethereum ; https://www.coindesk.com/tech/2020/12/21/linzhi-begins-rollout-of-long-awaited-ethereum-miner-phoenix ; https://bitcoinmagazine.com/business/new-mining-manufacturer-linzhi-announces-ethereum-asic-miner ; https://linzhi.io/ ; https://github.com/Souptacular/linzhi |
| Jasminer X4 (Sunlune) | chip announced 6 June 2021; X4 server announced 11 October 2021, first batch shipped 29 October 2021; launch price not found (USD 497 used today) | **the only Ethash chip that moved the memory**: TechInsights found "the first ever DRAM-to-Logic hybrid-bonding" (wafer-to-wafer DBI), DRAM dies bonded face to face onto a 32 mm by 21 mm logic die on XMC's planar 40/45 nm node; Jasminer's words: "3DIC technology, by integrating the data storage unit and the computing unit on the same chip"; 5 GB per unit, 40 chips per X4 server; the DRAM vendor, capacity per die and node not found | 2,500 at 1,200 W: 2.08 (the X4-1U 520 at 240 W: 2.17) | 5.1x against the tuned 3090; a physics check: 64 reads of 128 bytes is 65,536 bits a hash, so 2.08 MH/J is about 7.3 pJ per bit all in, under any off-package DRAM | https://www.techinsights.com/ko/node/51986 ; https://www.techinsights.com/ko/node/52149 ; https://semiconductor-digest.com/?p=22931 ; https://miningnow.com/asic-miner/jasminer-x4-2500mh-s/ |
| Bitmain Antminer E9, E9 Pro | teased 27 April 2021 as "3 GH/s, the work of 32 GPUs"; shipped July 2022 at 2.4 GH/s, eight weeks before the Merge; E9 Pro February 2023, Classic only | conventional off-chip DRAM at larger scale: E9 (model 240-E) 6 GB in six bins 2,100 to 2,400 MH/s; E9 Pro (260-E) 7 GB; memory type not found on any page read; no teardown | E9 2,400 at 1,920 W: 1.25; E9 Pro 3,680 at 2,200: 1.67 | 3.0x and 4.1x against the tuned 3090 | https://d-central.tech/antminer-e9-family/ ; https://www.asicminervalue.com/miners/bitmain/antminer-e9-2-4gh ; https://www.coindesk.com/tech/2021/04/27/bitmain-to-release-antminer-e9-asic-for-ethereum-mining |
**Why the cap sits at 2x to 5x.** Rao's audit (6 September 2019): every DAG read is random at about 40 ns of latency "completely independent of the memory bandwidth or the computation engine"; "typical DRAM energy dissipation is 3 to 4 pJ per bit" and "the energy expended to move data from DRAM to compute are the same for GPU or ASIC", so the shipping chips showed "about 1.6x hashrate per watt over GPUs" (E3 0.24, A10 0.57, the best overclocked GPU 0.40 MH/W in his table); integrating memory with logic cuts the movement energy "much more than 10x" to "under 0.3 pJ per bit", "the looming threat" (https://github.com/ethcatherders/progpow-audit, the PDF's text, read 8 October 2026). Ren and Devadas (TCC 2017) give the bound: memory hardness bounds area, not energy; the energy of a memory access is comparable on a chip and a CPU, so bandwidth hardness is the only energy lever (https://eprint.iacr.org/2017/225). The derived ladder at 65,536 bits a hash: pure memory energy caps Ethash at about 0.76 MH/J on DDR3, 2.8 on GDDR6 and 3.9 on HBM2 (O'Connor et al., MICRO 2017: HBM2 3.92 to 3.97 pJ per bit, GDDR5 14.0); the shipped chips sit at 0.24 (E3), 0.6 (Innosilicon), 0.9 (Linzhi), 1.25 to 1.67 (E9, E9 Pro) and 2.1 (Jasminer, by leaving commodity packaging). The cap was the memory's own energy per bit, and the one chip that beat it moved the memory onto the die's face.
**The timeline.** Hash live July 2015; the E3 at 32 months (announced) and 36 (shipped); the first chip over 2x at 65 months (Linzhi, December 2020); 5x at 75 months (Jasminer, October 2021); the Merge at 86 (15 September 2022). The E3's death by DAG growth: Classic first, at epoch 328 (DAG about 3.56 GB, March 2020), then Ethereum, with a 30 March 2020 firmware stretching the DDR to about block 11.4 million (about October 2020): 20 to 27 months after shipping. Ethereum's responses: Zamfir's April 2018 poll (57 percent for an anti-chip fork); EIP-1057 created 2 May 2018, a 93 percent community vote in April 2019, audits delivered September 2019, "accepted" on 21 February 2020, EIP-2538's opposition on 25 February, then stagnant; the share claim in the EIP's own text: "as much as 40 percent of the Ethereum network may now be secured by ASICs" (undated inside a 2018 to 2020 document; no year-by-year series exists). After the Merge: Classic's hashrate went 64 to 183 TH/s in one day; today Classic reads 129.9 TH/s and ETHW 2.15; at USD 0.10 per kWh every Ethash chip in the table loses money (E9 Pro minus USD 5.28 a day), and a later wave (iPollo V1 3.6 GH/s at 3,100 W, June 2022; Jasminer X16-P 5.8 GH/s at 1,900 W, August 2023) holds Classic (https://hashrateindex.com/blog/how-much-ethereum-mining-hashrate-can-other-blockchains-absorb/ ; https://2miners.com/etc-network-hashrate ; read 8 October 2026).
**The mapping to v6.** Ethash is class C in full, and its chips are the f = 1 chip of `chip-model-v3.md` section 5 at three points on the packaging ladder: commodity DRAM on a board (E3, E9: 1x to 4x), memory sized and placed per chip (Linzhi: 2x), DRAM bonded to the logic (Jasminer: 5x, the model's HBM3 row). None of v6's four layers touches a chip of this class: the program, the mixer, the op mix, the family schedule and the acceptance floor are all firmware or configuration to a controller that stores the dataset; the only layer that reaches it is layer 2, and only when the dataset passes the chip's board, which at 2 GiB plus 0.5 GiB a year is year 60 for a 32 GB board (section 4.2). Does v6 close it: **no**. The number: 5.1x per joule on GDDR7 and 7.5x on one HBM3 stack at zero premium against the 5090 unlocked, 3.6x at the 5090's 1,300 MHz knee, 2.1x at the knee with the class v4 shadow at k = 1 (`counter-asic-4-research.md` section 0); the history's measured band for exactly this chip class is 1.0x (E3) to 5.1x (Jasminer), and Jasminer's number is the model's HBM-class row reached in 2021 on a 40 nm logic die. The forward line of the same mechanism is lane B's (`docs/analysis/class-v6/hardware-future.md`, master 34f63b3c, section 4.4 and its table row for fine-grained hybrid-bonded DRAM on logic: 13x to 17x per joule at zero shadow, modelled, on the 2028 to 2030 roadmaps); the Jasminer X4 is that row's shipped precedent, five years early and at 5x on a planar node, which is the reason to read lane B's 13x to 17x as a ceiling a first product will not reach and a second one will approach. The one thing Igneum has that Ethash did not: the honest card is latency-bound at 4-byte reads, not bandwidth-bound at 128, so the chip's energy per read is the activate's 909 pJ plus a 32-byte atom (2.0 nJ on GDDR7 against the card's measured 8.7 to 10.9 nJ marginal), which is where the 5.1x comes from, and the shadow is the only term on the card's side of that ratio.
### 2.3 RandomX (Monero, 30 November 2019): the chips, RandomX v2, and the X9's withdrawal
**The mechanism of the hash.** A VM running 8 chained programs of 256 instructions, 2,048 iterations each, over a 2 MiB scratchpad in three tiers (16 KiB, 256 KiB, 2 MiB) and a 2,080 MiB dataset derived from a 256 MiB cache by SuperscalarHash, a random superscalar program of about 450 instructions with 155 64-bit multiplies per function, tuned to a 170-cycle latency to match DRAM; double-precision floating point in all four rounding modes; the light-mode chip (cache on die) pays 760 cycles and 1,240 multiplies per item, "energy comparable to loading 64 bytes from DRAM" (https://github.com/tevador/RandomX/blob/master/doc/design.md and doc/specs.md, read 8 October 2026). Its DRAM argument: "DRAM cannot do more than about 25 million random accesses per second per bank group", about 1,500 H/s per bank group.
**The chips, and what each specialised.**
| Chip | Date | What is known of the inside | Rate, watts | Per joule against the best CPU | Source (read 8 October 2026) |
|---|---|---|---|---|---|
| Bitmain Antminer X5 | announced 27 August 2023, shipped September 2023 | "Bitmain's first RISC-V architecture CPU" (the reseller's only line); Spagni: not an ASIC but "a board containing multiple RISC-V CPU chips"; SChernykh: the chips were likely in use from about 2021, two years before sale, and do not beat Ryzen rigs per joule; core model, count, node, DRAM type and amount: NOT FOUND, no teardown | 212 kH/s at 1,350 W (157 H/W) | 1.46x over a Ryzen 9 7950X (107.5 H/W on Kryptex); a 100 W-capped Ryzen 9 9950X at 199 H/W beats it | https://criptonoticias.com/mineria/bitmain-lanza-antminer-x5-mineria-monero-asic ; https://bt-miners.com/products/bitmain-antminer-x5-monero-miner-212k-bt-miners/ ; https://pool.kryptex.com/en/device/cpu/amd/ryzen-9-7950x |
| Bitmain Antminer X9 | sales opened 26 December 2025 at USD 5,600, shipping scheduled July 2026; WITHDRAWN by mid-May 2026, refunds within hours, zero units shipped, "technical adjustments and a new strategy" through resellers, no Bitmain statement | "custom RISC-V cores specifically optimized for RandomX" (Bitmain's claim as relayed in Monero issue 10270); nothing else | 1,000 kH/s at 2,472 W (404 H/W), claimed | about 3.8x over the 7950X, claimed, never measured | https://bitmain.com.vc/news/bitmain-launches-antminer-x9 ; https://github.com/monero-project/monero/issues/10270 ; https://oneminers.com/blogs/news/whatever-happened-to-the-antminer-x9-bitmain-monero-miner (2 October 2026) |
| Pinecone INIBOX R1X | launched March 2026; shipping windows slipped from August to 10 to 18 October 2026; USD 3,200 to 4,950 | "built from ground up silicon", "optimized memory architecture for RandomX"; cores, node, DRAM: NOT FOUND; no delivered unit tested | 1,200 kH/s at 2,055 W (584 H/W), claimed | about 5.4x over the 7950X, claimed, undelivered | https://pineconebox.com/product/3 ; https://millionminer.com/news/monero-mining-guide-2026-antminer-x5-x9-pinecone-r1x-randomx |
| tevador's own "possible ASIC design" (24 December 2018, pre-release, marked outdated) | | 4 GiB of HBM for the dataset, 64 MiB of SRAM for 256 parallel 256 KiB scratchpads, 256 decoder and scheduler cores, 28 single-instruction workers; about 120,000 programs a second at about 300 W, "9 times more power efficient than a CPU" | | 9x, by the designer's own estimate of the pre-release design | https://github.com/tevador/RandomX/issues/11 |
**What the design missed, in its authors' words.** RandomX v2 (PR 317 by SChernykh, merged 17 February 2026; v2.0 released 25 March 2026; Monero mainnet activation in PR 10038, open since August 2025, no date) names the three gaps a chip or a "specially designed CPU" took: (1) CFROUND, the rounding-mode switch, "costs up to 10 percent of hashrate on Ryzen CPUs" and "this is where an ASIC or a specially designed CPU can get an easy advantage"; v2 switches rounding 16 times less often; (2) the scratchpad initialisation was the only AES, so "a dedicated circuit for scratchpad initialization" paid off; v2 puts 16 AES operations per iteration in the main loop; (3) "while CPU cores got faster over the years, RAM latency stayed basically the same", about 50 to 55 ns from tuned DDR4 in 2019 to tuned DDR5 in 2026, so the fixed 256-instruction program left latency slack a faster core could not fill; v2 lengthens the program to 384 and prefetches two iterations ahead. Work per hash rises 52.9 percent; measured CPU hash rates move from minus 12.9 percent (a 100 W-capped 9950X) to plus 8 percent (a 28 W laptop part) (https://github.com/tevador/RandomX/blob/master/doc/design_v2.md and /pull/317, read 8 October 2026). The press reading: v2 "doesn't seem to be an attempt to eliminate every form of specialization", it removes "unintended advantages that benefited hardware in version 1.0" (https://www.coinpro.ch/en/?p=42081, 31 March 2026).
**The timeline.** Live 30 November 2019; X5 at parity hardware 46 months later (September 2023), on chips SChernykh believes mined privately from about 2021 (21 months after launch); the X9 announced at 73 months and withdrawn at 78; v2 released at 76 months with a 52.9 percent work increase and no activation date; the R1X undelivered at 82 months. The X9's withdrawal is read by the trade press as Bitmain waiting for v2 rather than shipping a part the fork would hit (https://oneminers.com/blogs/news/antminer-x9-cancelled-what-bitmain-pulling-the-model-means-for-monero-mining, 31 July 2026). Corrections to the 5 October file's row 17: the X9 was not delivered in July 2026 (withdrawn, zero units), and "no fork as of October 2026" is now "v2 released, activation pending".
**The mapping to v6.** RandomX is class D (the firmware-survivable machine): every X5 claim is a many-core RISC-V board over commodity DRAM, a better CPU for a fixed VM spec, and the v2 fixes are exactly the parameters class v6 layer 1 draws (program length, the op mix's cost on the honest machine, the memory latency slack). Two readings bind: first, RandomX's whole gap is the CPU's out-of-order overhead against an in-order many-core board, and Igneum's honest machine is already the in-order many-lane design, so the X5's 1.46x has no Igneum analogue; the Igneum analogue of "a better machine for the fixed spec" is class C's memory chip at 5.1x. Second, the v2 changes show what a per-era draw of program length and op mix buys: it closes the slack a faster honest core leaves (the latency-shadow argument of class v4 in RandomX's words), and it costs the honest machine up to 12.9 percent of rate at the power-capped point, which is the Igneum premium question in another chain's numbers. Does v6 close it: yes for the class (layers 1 and 3 draw what v2 had to fork to change), and the number is the X5's 1.46x, which becomes the shadow's premium arithmetic on Igneum (2.1x at k = 1 at the knee).
**The Qubic episode (2025) as the detector's lesson.** Qubic's pool reached an average of 22.09 percent of Monero's hashrate over the campaign and 23 to 34 percent during ten withholding periods, with six-hour windows near 50 percent and never a daily 51 percent (Lee and Kim, arXiv 2512.01437, AFT 2026, read 8 October 2026); an 18-block reorg on 14 September 2025 invalidated 117 to 118 transactions. Detection rested on things the adversary controlled: one payout wallet, extra-nonce signatures in its coinbase, its own pool API; when Qubic encrypted its job messages and rotated keys those signals went, and Rucknium's warning stands that a miner split across addresses and solo-mining leaves only the orphan rate and double spends as signals. Qubic ran stock CPU miners, so the nonce-distribution method that found the 2018 and 2019 chips (MoneroCrusher: nonces clustered under about 1.34 billion of 4.3; 85.2 percent of the hashrate, about 5,400 machines at 128 kH/s) did not apply. For Igneum's detector (the 5 October addition 4) this means two instruments, not one: the per-program rate spread and nonce pattern for a chip, and a share-by-key-and-template pattern for a concentrated honest fleet; neither survives an adversary who randomises both.
### 2.2 ProgPoW (EIP-1057, May 2018): the independent review, the exploit, and the adopters
**The mechanism.** A random program re-drawn every PROGPOW_PERIOD (50 blocks in 0.9.2, 10 blocks, about 2 minutes, in 0.9.3) from the block number, so miners compile ahead; 16 lanes, a 32-register file per lane, 64 outer iterations each with 4 uint32 DAG loads per lane (256 bytes per lane-group read), 11 cache accesses into a 16 KB cache and 18 random math ops drawn from eleven (add, mul, mulhi, min, rotl, rotr, and, or, xor, clz, popcount) with KISS99 as the generator and FNV1a for merging; keccak-f800 with 32-bit words at both ends "to reduce impact on total power"; the stated aim is that "the algorithm's requirements match what is available on commodity GPUs", the stated chip gain "minimal, roughly 1.1 to 1.2x", with the remaining chip levers named as removing the graphics pipeline, the floating-point units and minor merge-function tweaks (https://eips.ethereum.org/EIPS/eip-1057 and https://github.com/ifdefelse/ProgPOW, read 8 October 2026).
**The independent review, exactly.** Least Authority (report version 9 September 2019): no issues, five suggestions. Suggestion 2, the light-evaluation attack, in the report's words: "on-die scratchpad memory of around 100 MB is possible in ASICs, that we can fetch at least 128 bytes during a single read, and that such a read might have a latency in the range of one to some tens of cycles"; a chip replaces every DAG read with calc_dataset_item(cache, i) over an on-die cache, at a latency near DATASET_PARENTS x k1, "as low as about 300 cycles"; "the energy expended per bit to access DRAM is about 3 pJ per bit, but when the memory access is on-chip, it decreases to 0.3 pJ per bit, which is a 10x improvement"; conclusion: "efficient light-evaluation attacks may become possible within a few years. This is also an issue that applies to Ethash"; the mitigation offered: raise DATASET_PARENTS from 256 to 512 (which 0.9.4 did), and "for details on the related hardware advancements, please see Bob Rao's corresponding audit report". Suggestion 5: "hardware targeting machine learning is also useful for ProgPoW mining" (the PDF at https://leastauthority.com/static/publications/LeastAuthority-ProgPow-Algorithm-Final-Audit-Report.pdf, text extracted, read 8 October 2026). Bob Rao (6 September 2019): "the only meaningful metric is Energy per Hash"; Ethash chips "about 1.6x hashrate per watt over GPUs"; "10/7 nm processes provide up to 25 Mbits per mm^2 of SRAM and 100 M transistors per mm^2"; "with sufficient on-chip memory available, ProgPOW ASICs with << 0.1x E/H over GPUs can be built"; three chip approaches costed (the whole DAG on die, "possible in 2025+"; a custom stacked memory; the cache only, "about 51 MB as of 30 August 2019", with the item recomputed, "512 MB SRAM" on the slide); the DAG-on-die economics on a three-year Moore cadence: a single die holding a 6.49 GB DAG in 2024 to 2025 at 532 mm^2 and USD 221 per good die (USD 34 per GB), or sixteen dies of 33 mm^2 at USD 6.51 each; "a die that can hold the logic and entire DAG at any point in time becomes cost effective at around 2025 and beyond"; the 16-die split "becomes cost effective today": 16 x USD 6.62 of silicon plus USD 16 of package plus USD 25 of PCB plus USD 25 of heatsink and interface, "about USD 172 total", against "about USD 240" for a GPU board with "8 GB GDDR6 about USD 150"; a 10 nm-class chip "USD 20 M+" and "1+ year to develop, can be done if there is a 150-day ROI to miners" (the PDF at https://github.com/ethcatherders/progpow-audit, text extracted, read 8 October 2026).
**The exploit.** Kik, 4 March 2020: the 64-bit seed carried between the two keccak passes is too small; fix a seed and compute its mix once, grind an extra-nonce in the header to meet the difficulty on the final keccak, then scan nonces until keccak_progpow_64(header_hash, nonce) equals the seed; the memory path runs once per 2^64 nonces and the rest is keccak, "ASICs benefit most when network difficulty exceeds 2^50"; 0.9.4 widened the carried state from 64 to 256 bits (the digest of the first keccak plus the mix plus padding) (https://github.com/kik/progpow-exploit ; https://github.com/ifdefelse/ProgPOW ; read 8 October 2026).
**Linzhi's claim.** 8 January 2019: "shocked" by ProgPoW with USD 4 M invested, and a stated intention "to study the feasibility, and then build, ProgPoW ASICs"; the repository recording their claim puts a ProgPoW chip at 3x to 8x (https://github.com/Souptacular/linzhi ; https://forklog.com/proizvoditel-majnerov-linzhi-vystupil-protiv-realizatsii-predlozheniya-progpow/ ; read 8 October 2026). No ProgPoW chip was ever shown.
**The timeline and the adopters.** EIP created 2 May 2018; a 93 percent vote of 2.93 M ETH in April 2019; both audits September 2019; "accepted" on the 21 February 2020 call; EIP-2538's opposition 25 February; the 6 March 2020 call with "frustration but little progress"; stagnant since; the Merge 15 September 2022. KawPow (Ravencoin, 6 May 2020), FiroPoW (26 October 2021), ProgPowZ (Zano), Sero and Quai (January 2025) run ProgPoW variants; no chip is listed for any of them on WhatToMine or asicminervalue as of 8 October 2026 (https://whattomine.com/coins/234-rvn-kawpow/gpus ; https://www.asicminervalue.com/), on prizes that never reached the market caps at which the 2018 chips appeared (the 5 October file, section 2.5).
**The mapping to v6.** ProgPoW is class D (a sequencer over eleven ops reads the per-period program as configuration) and class E (Kik). Its review is the one piece of the record that priced Igneum's own chips before Igneum did: the light-evaluation attack is the f = 0 recompute chip (M16, 0.92x at the op budget with the mixer at x8), and Rao's 16-die DAG holder at USD 172 is the f = 1 chip at USD 470 of memory and board. Does v6 close it: the class D half yes (layers 1 and 3 are ProgPoW's period change made an era draw, with no fork), the class E half yes (Igneum's seed is 256 bits through the VDF; layer 4 is the acceptance-side test the audits said to add), and the Rao half no (class C, section 2.1). The number: ProgPoW claimed 1.1x to 1.2x against a conventional compute chip and never priced the memory-system chip; Igneum's model gives that chip 5.1x.
### 2.4 Cuckoo Cycle (Grin, January 2019): the GRN1, the G32, the iPollo G1 and the 2025 bounty
**The mechanism of the hash.** Find a 42-cycle in a random bipartite graph of 2^31 or 2^32 edges from siphash; the lean solver keeps "1 bit per edge and 1 bit per node in one partition", bottlenecked by random node-bit access, which "requires tons of SRAM, which is lacking on CPUs and GPUs, but easily implemented in ASICs"; the mean solver keeps 33 bits per edge and is bandwidth-bound, about 4x faster; Tromp: "our primary PoW of Cuckatoo31+ is intended to be mined by ASICs"; the family is "a Proof of SRAM" (https://github.com/tromp/cuckoo ; https://forum.grin.mw/t/cuckatoo31-im-mutability/2442 ; read 8 October 2026). The memory geometry a chip needs (the 13 November 2018 feasibility thread): the 512 MB edge bitmap is accessed sequentially and can sit in external DRAM (96 GB/s with 128 MB on chip); the node bitmap must be SRAM; Cuckatoo31 fits one die at "256 + 64 = 320 MB of on-chip memory", Cuckatoo32 needs "at least 640 MB" or 512 MB of SRAM plus 512 MB of serial DRAM; "trimming is over 99 percent of the effort" (https://forum.grin.mw/t/cuckatoo32-feasibility/1199 ; https://forum.grin.mw/t/advice-on-cuckatoo-hardware-implementation/12042). The original "several orders of magnitude" time-memory claim fell to Andersen's edge trimming on 31 March 2014, two months after publication, and the design took it as its baseline.
**The chips.**
| Chip | Dates | Mechanism | Rate, watts, price | Against a GPU | Source (read 8 October 2026) |
|---|---|---|---|---|---|
| Obelisk GRN1 (Cuckatoo31) | announced 17 January 2019; chip details 20 March; sale 9 April (Mini 70 GPS at 400 W for USD 2,000; GRN1 420 GPS at 2,200 W for USD 10,000; Immersion 840 at 4,400 W for USD 20,000; shipping October 2019); CANCELLED 19 July 2019 with full refunds | one die, TSMC 16 nm, "a full 512 MiB of memory on board" (SRAM), 150 GPS at 800 W per chip, "thousands of hashing cores and memory banks", siphash plus blake2b, two sorter types; cancelled for the Cuckatoo31 phase-out, Grin under USD 2 in May 2019 and funding | 420 GPS at 2,200 W planned | about 50x per joule against a GTX 1080 Ti at about 0.9 GPS and 250 W on C31 (approximate), planned, never built | https://forum.grin.mw/t/obelisk-grn1-chip-details/4571 ; https://forum.grin.mw/t/obelisk-grn1-full-sale/4773 ; https://forum.grin.mw/t/grn1-cancellation-announcement/5624 |
| Innosilicon G32 (C31+ and C32+) | announced 17 April 2019 (G32-Mini 21.5 GPS at 140 W for USD 788; G32-1800 328 GPS at 1,800 W for USD 9,388; delivery from August 2019); CANCELLED 16 January 2020 citing foundry delays | about 100 chips at about 1.8 to 1.9 GPS each on C32 (an investor's figure) | never shipped | | https://forum.grin.mw/t/innosilicon-grin-miner-g32-preliminary-specification/4842 ; https://forum.grin.mw/t/innosilicons-grin-asics-canceled/6932 |
| iPollo G1 (Cuckatoo32) | December 2020 | 30 chips at 12 nm; whether the node bits sit in SRAM or the edge bitmap in DRAM is not published; Tromp doubts it is multi-chip in Innosilicon's sense | 36 GPS at 2,800 W, USD 9,000 | about 3x per joule against an RTX 4060 at 0.45 GPS and 110 W | https://pool.kryptex.com/device/asic/ipollo/g1 ; https://miningboard.com/algorithms/Cuckatoo32 |
**The 2025 bounty.** Stephan Theisgen claimed the USD 10,000 linear time-memory trade-off bounty on 11 April 2025 (a solver at N/k bits at most 10k times slower, any k at or above 2), paid 30 April; the measured penalty about half an order of magnitude past linear; Tromp's reading: a chip with N/k bits must hash each edge "roughly (k + 1000) times" against under 6 for the lean miner, so memory can be traded for time but not for energy, and Cuckatoo "remains a Proof of SRAM" (https://forum.grin.mw/t/another-cuckatoo-bounty-succesfully-claimed/11739, read 8 October 2026).
**The mapping to v6.** Class B. Cuckoo's resource was SRAM size at 320 to 640 MB per die, and a 16 nm die with 512 MiB of SRAM was designed, priced and sold before the economics killed it; the one chip that shipped sat at 3x because half its working set stayed in DRAM. Layer 2 answers it the way the chip model already does: Igneum's 256 MiB cache is the GRN1's die (128 mm^2 at N5, `sram-mirror.md`), and the 2 GiB dataset above it is what the lean solver never had to hold; the time-memory curve Cuckoo got wrong by 50x and then bounded in 2025 (linear in time, not in energy) is the curve `chip-model-v3.md` section 5 draws for Igneum (monotone; the f = 0 end pays 6.3 nJ and 9,360 ops per item against 1.2 to 2.0 nJ for a stored one), and its verdict is the same as Tromp's: memory can be traded for time, not for energy. Does v6 close it: yes, and it was closed at class v3. The number: 50x planned on C31 and 3x shipped on C32; Igneum's f = 0 chip 1.86x per joule and the f = 1 chip, which Cuckoo did not have because its memory was never a DRAM-scale random-read set, 5.1x.
### 2.5 Equihash 200,9 (Zcash, October 2016): the Z9 and the parameter-following chip
**The mechanism of the hash and its claim.** Wagner's generalised birthday problem with algorithm binding; the paper's claim: a PoW needing "700 MB of RAM" that "increases the computations by the factor of 1000 if memory is halved" (https://eprint.iacr.org/2015/946, read 8 October 2026). Zcash chose (200, 9), which solvers run in about 144 MB (approximate); no trade-off-resistance bound was ever proved for Equihash (Alcock and Ren, 2017, cited through the Fudan paper below).
**The chips and the mechanism.** Bitmain's BM1740 (Z9 mini, announced 3 May 2018 at 10 kSol/s and 300 W, shipped June; Z9 September 2018, 42 kSol/s at 970 W on 48 chips) is "a single-chip Equihash miner", "presumably with around 128 MB of memory", eDRAM or SRAM "an open question" (Tromp, 7 June 2018, https://forum.z.cash/t/let-s-talk-about-asic-mining/27353/3332): the whole (200, 9) working set on one die, which the 1,000x claim had assumed impossible at that size. A (144, 5) solver needs over 1.6 GB and cannot fit; Tromp "physically inspected the product and did not find enough memory to handle (144, 5)". Vorick's architecture (13 May 2018) is the other route: "a basic architecture for equihash ASICs that would be able to successfully follow a hardfork that chose any set of parameters", with "massive speedups and efficiency gains over GPUs", because on a chip "you can merge the memory and computation together ... do most of your manipulating in-place" (the essay, through https://steemit.com/crypto/@waraa/the-state-of-cryptocurrency-mining). The published design of that route: Bai, Gao, Hu and Zhang, NDSS 2019, an adversary solver whose sort step is a linear-time insertion sorter of 2,048 "smartcell" flip-flop chains feeding merge stages buffered in off-chip DDR4 (list 1,600 Mib, pairs 2,016 Mib for (200, 9)), pair generation and XOR on small MCUs; simulated at SMIC 28 nm: 40.6 Sol/s at 500 MHz for 0.75 to 0.78 W, 52 to 54 Sol/J against about 4 for the best GPU software, "at least 10x" and parameter-independent (https://www.ndss-symposium.org/wp-content/uploads/2019/02/ndss2019_09-5_Bai_paper.pdf, read 8 October 2026). Later chips: Innosilicon A9 (June 2018, 50 kSol/s at 620 W, USD 9,999); Z11 (April 2019, 135 kSol/s at 1,418 W, 12 nm); Z15 (420 kSol/s at 1,510 W); Z15 Pro (June 2023, 840 kSol/s at 2,780 W). Against a GTX 1080 Ti at 735 to 785 Sol/s and 250 to 305 W (about 2.7 Sol/J): Z9 mini 12x, Z9 16x, A9 30x, Z11 35x, Z15 Pro 110x (https://www.asicminervalue.com/miners/bitmain/antminer-z9 ; https://www.asicminervalue.com/miners/bitmain/antminer-z15-pro ; https://en.wikibooks.org/wiki/ZCash_mining_GPU_Comparison/GPU_Mining ; read 8 October 2026).
**The timeline.** Hash live 28 October 2016; three groups mining on secret chips before the Z9 announcement (Vorick); the Z9 mini at 18 months; the Zcash Foundation's statement of 8 May 2018 asked whether chips "could handle different parameters of Equihash" and the community vote of June 2018 went 45 to 19 against prioritising resistance; no fork; proof of stake announced November 2021; the Z15 Pro at 80 months.
**The mapping to v6.** Class B, with a lesson for layer 1: Equihash's parameters (n, k) were the knob the forks turned (Bitcoin Gold to 144,5; Beam to 150,5; Flux to 125,4), and a chip was designed to follow "any set of parameters" by keeping the memory off die and the sort on it. A per-era draw of a parameter a chip can follow is a configuration to it; a draw of the memory SIZE is the one that forced the Z9's single die to fail on (144, 5), which is layer 2's mechanism, not layer 1's. Does v6 close it: yes by layer 2 (a 2 GiB set is not a die), and the parameter-following chip is the honest warning for layer 1's draws. The number: 12x at the first chip, 110x by 2023, against a hash whose 1,000x penalty claim never had a proof; Igneum's claim for its curve rests on a drawn curve (`chip-model-v3.md` 5.4) and the in-house pass's exact pebbling optimum (adv-cache-3), and still has no proof.
### 2.6 Scrypt and Argon2: the Litecoin chips, the lookup gap, and MTP
**Scrypt (Tenebrix and Litecoin, 2011; N = 1,024, r = 1, p = 1: a 128 KB scratchpad).** The mechanism of every scrypt chip is one 128 KB of on-die SRAM per hashing core and many cores per die: Watkins (2014) "there only needs to be 128 KB of memory per processor ... one core driving a 128 KB cache", SRAM chosen over every other memory as fastest, salsa20/8 about 60 percent of the runtime and memory access 38 percent (https://arxiv.org/pdf/2208.02160); the BM1485 of the L3+ (June 2017, 504 MH/s at 800 W): 12 cores per chip, "every BM1485 integrates on-die SRAM to hold that scratchpad", 28 nm, 288 chips per unit, about 1.7 MH/s per chip (so about 1.5 MB of SRAM per chip, arithmetic) (https://d-central.tech/mining-glossary/bm1485/ ; https://www.asicminervalue.com/miners/bitmain/antminer-l3-504mh); the L7 (November 2021, 9.5 GH/s at 3,425 W, USD 15,000, 0.36 W per MH against the L3+'s 1.58): the BM1489 at TSMC 7 nm, 480 chips (https://d-central.tech/mining-glossary/bm1489/ ; https://cryptoage.com/en/2550-bitmain-antminer-l7-is-a-new-asic-miner-for-litecoin-and-dogecoin.html); the Innosilicon A2 (21 April 2014, 28 nm, 1.6 to 1.8 MH/s per chip at 10 W, about 150 MH/s per box at 1 kW: https://www.design-reuse.com/news/34403/innosilicon-28nm-litecoin-asic-reference-miner.html); the KnC Titan (March 2014, 250 MH/s at 800 to 1,000 W, USD 9,995, "4 chips x 2,284 cores": https://www.coindesk.com/markets/2014/03/28/kncminer-updates-titan-spec-promises-250mhs/), whose 2,284 cores cannot each hold 128 KB on a 2014 die, so it shared scratchpads or took the time-memory trade-off (not confirmed). The trade-off itself is in the designer's record: Percival (18 November 2012) on storing every other scratchpad entry: memory halved for about 25 percent more BlockMix work, area-time about 0.625x, the trade favouring the attacker at 2x to 4x reductions, "already in the paper's cost estimates" (https://mail.tarsnap.com/scrypt/msg00092.html; all read 8 October 2026). Against an R9 280X at 700 to 740 kH/s and 340 to 450 W at the wall: A2 about 70x, L3+ about 300x, L7 about 1,300x per joule (derived, approximate). Scrypt-N (Vertcoin 2014: N doubling by timestamp up to 2^30) was dropped on 13 December 2014 for Lyra2RE "as a proactive defense against emerging Scrypt-N capable ASICs", because raising N "simply involves doing more iterations" and more SRAM or a larger lookup gap (https://vertcoinproject.org/vertcoin_whitepaper.pdf ; https://coincentral.com/what-is-vertcoin-a-beginners-guide/).
**Argon2 as a proof of work (MTP, Zcoin, 10 December 2018).** Argon2d over 4 GB with a Merkle tree; Dinur and Nadler (2017): malicious proofs with under 1 MB, 1/3,000 of the honest memory, at a computation penalty of 170, "more than 55,000 times faster than what is claimed by the designers", with a 2^64 one-time precomputation, by injecting chosen blocks that steer Argon2d's data-dependent addresses (https://eprint.iacr.org/2017/497); MTP 1.2 patched it before launch; no Argon2 chip was ever built; Firo replaced MTP with FiroPoW on 26 October 2021 (https://firo.org/2021/10/01/firopow-and-instantsend-release.html). Yescrypt and yespower (GlobalBoost-Y 2014; Yenten, Cranepay, Tidecoin on yespower from 2018; L2-latency-bound sequential work: https://www.openwall.com/yespower/): no chip found, small prizes.
**The mapping to v6.** Class B throughout, and the parameter-growth lesson for layer 2 in its oldest form: Scrypt-N's automatic growth was abandoned because a chip follows a scratchpad that grows by doubling the SRAM it already has, and Percival's own 2012 note says the time-memory trade favours the chip at 2x to 4x. Does v6 close it: yes by layer 2 (no die holds 2 GiB; the curve is monotone against partial stores); the lookup-gap lesson is the reason the dataset's chained cache has a drawn pebbling optimum (adv-cache-3) rather than a claim. The number: 1,300x for scrypt by 2021; Igneum's f = 0 chip at 1.86x per joule. MTP is class E (section 4.4).
### 2.7 to 2.9 KawPow, Autolykos, Octopus: the no-chip hashes, and why
| Hash | Live | Mechanism | Chip status, 8 October 2026 | Why no chip (the honest reading) | Source |
|---|---|---|---|---|---|
| KawPow (Ravencoin; Neoxa, Clore, Meowcoin, Neurai) | 6 May 2020 | ProgPoW 0.9.4 with a per-block program; "no additional future algorithm forks are envisaged" | none listed on WhatToMine or asicminervalue | class D on a small prize: Ravencoin's cap never reached the 2018 cluster's | https://whattomine.com/coins/234-rvn-kawpow/gpus ; https://github.com/RavenProject/Ravencoin/blob/master/roadmap/README.md |
| Autolykos v2 (Ergo) | February 2021 (v1 July 2019) | a k-sum (k = 32) over a Blake2b table of 2^26 elements of 31 bytes (2.08 GB) growing about 5 percent per 51,200 blocks from block 614,400 to a cap of 2,143,944,600 elements at block 4,198,400; v1's non-outsourceable puzzle removed because "large players could bypass this resistance using smart contracts" | none | class C territory (a table read per hash) on a small prize; its growth rule is the one automatic schedule in the record untested by a chip; f2pool and Ergo's own docs call it resistant with no chip named | https://docs.ergoplatform.com/mining/autolykos/ |
| Octopus (Conflux) | October 2020 | Ethash-style DAG (the "dense matrix step" unverified in the source) | none | as above; f2pool: "cannot be efficiently mined with FPGAs or ASICs" | https://f2pool.io/mining/guides/how-to-mine-conflux/ |
| Verthash (Vertcoin), FishHash (Iron Fish from April 2024, Karlsen) | January 2021; April 2024 | a 1.2 GB table from the chain's headers; a 4,608 MB constant dataset with Blake3 and 512 iterations | none | the same class as Ethash's chips, on prizes under the 2018 cluster | https://fips.ironfish.network/fips/fip-3-memory-hard-mining-algorithm |
The reading for v6: "no chip" on a DRAM-scale random-read hash is an economic fact, not a design one (the 5 October file, section 2.5: compute-bound hashes got chips at USD 20 K to 30 K of daily issuance, Ethash at USD 7.6 M). Every hash in this table is class C and none of them has been tested at a prize that pays for a controller project; the Ethash record is the test, and it read 1x to 5x.
### 2.10 kHeavyHash (Kaspa, November 2021): the KS chips
**The mechanism.** cSHAKE256 of the header and nonce, a 64 x 64 matrix of 4-bit values generated from the pre-PoW hash, a nibble-wise matrix-vector multiply, XOR, a final cSHAKE (https://github.com/kaspanet/rusty-kaspa/blob/master/consensus/pow/src/lib.rs, read 8 October 2026); designed for optical and specialised hardware; the chip wires the multiply as a fixed pipeline. IceRiver KS0 (September 2023, 100 GH/s at 65 W), KS1 (1 TH/s at 600 W), Antminer KS3 (August 2023, 9.4 TH/s at 3,550 W), KS5 Pro (March 2024, 21 TH/s at 3,150 W), KS5L (April 2024, 12 TH/s at 3,400 W); against an RTX 4090 at 2.08 GH/s and 226 W: KS0 about 170x, KS5 Pro about 720x per joule (https://www.asicminervalue.com/miners/bitmain/antminer-ks5-pro-21th ; https://www.kryptex.com/overclocking/nvidia-rtx-4090-micron-24gb-medium-overclock ; read 8 October 2026). Timeline: 17 to 20 months to the first chip; the GPU share negligible by late 2023; seven forks left Kaspa to re-resist (the 5 October file, row 23). Node: not found.
**The mapping to v6.** Class A. The matrix multiply is a warning for the reserve's mm8 family, not for the hash: a fixed 64 x 64 nibble multiply is the cheapest thing silicon does, and the measured rows agree (the 5090's int8 MAC at 1.4 to 4.1 pJ against a 5 nm array's claimed 0.04 to 0.4, `counter-asic-4-research.md` 15.1a). Does v6 close it: yes, closed since class v2 (no fixed function to unroll); the number, 720x, is what a fixed pipeline does to compute-bound work, and the reserve's ordering (section 4.3) keeps mm8 last for exactly this reason.
### 2.11 CryptoNight (Bytecoin 2012, Monero 2014): the X3, the secret chips, and four forks
**The mechanism of the hash.** 524,288 iterations of an AES round plus an 8-byte multiply over a 2 MB scratchpad sized to a 2014 per-core L3; latency-bound at SRAM scale.
**The chips.** Bitmain Antminer X3, announced 15 March 2018 at 220 kH/s and 550 W (465 to 470 W measured), 180 BM1700 chips on three boards, USD 11,999 for batch 1 falling to USD 1,900; no teardown or vendor description of the BM1700 exists (node, SRAM, AES units not found). The arithmetic is the mechanism: 1.2 kH/s and about 2.6 W per chip is one or two 2 MB scratchpads in on-die SRAM with a hardware AES round per chip, and nothing else reaches that rate in that power (approximate, inference). Baikal Giant N, March 2018, 20 kH/s at 60 W. Against a Vega 64 at 2,009 H/s and about 200 W card power (about 10 H/W) the X3 is about 40x per joule; against a Threadripper 1950X at about 1,000 H/s and 185 W, about 75x (approximate) (https://www.asicminervalue.com/miners/bitmain/antminer-x3-220kh ; https://www.asicminervalue.com/miners/baikal/bk-n ; https://hothardware.com/reviews/monero-mining-with-amd-ryzen-threadripper ; read 8 October 2026).
**The secret chips.** Monero's own 2018 review: by early 2018 "it was estimated that 80 to 90 percent of the network was specialized hardware" (https://web.getmonero.org/2019/02/12/2018-year-in-review.html); about half the hashrate (about 500 MH/s of 1,000) left at the v7 fork on 6 April 2018; Krawiec-Thayer's nonce study (24 November 2018) found half of all blocks with nonces in the lowest 0.002 percent of the space, patterns that "evaporated abruptly" at the fork (https://www.hackernoon.com/utter-noncesense-a-statistical-study-of-nonce-value-distribution-on-the-monero-blockchain-f13f673a0a0d). Vorick (13 May 2018, through secondary coverage): secret Monero ASIC mining "since early 2017, making up 50 percent of the hashrate". The timeline from the hash to the first secret chip is therefore about 33 months (April 2014 to early 2017), not the 43 the 5 October file gives (which counted to the fork); the announced chip came at 47.
**The four forks and what each cost a chip.** v7 (6 April 2018): a one-byte tweak to the main loop; half the hashrate left; "a precaution and deterrent". v8 (18 October 2018, height 1,685,555): a whole-cache-line shuffle (4x the bandwidth demand) plus a 64:32 division and a 64-bit square root per iteration, 5 to 10 percent off CPU rate; the hashrate went from about 320 MH/s to just under 1,000 MH/s before March 2019 with chip nonce patterns visible from December: chips back inside two months, dominant inside four (https://decrypt.co/14421/ ; https://en.cryptonomist.ch/2019/05/08/mining-monero-hashrate-pow-change/). CryptoNight-R (9 March 2019, brought forward from April after the detection): a per-block random sequence of 60 to 69 integer instructions (63 on average; MUL 40 percent, XOR 23, SUB 12, ADD 12, ROR 8, ROL 6) over 9 registers, seeded by height so miners compile ahead; SChernykh's claim is a chip's minimum latency for the random math "at least 2.5 times higher" than the DIV plus SQRT it replaced (a chain of 15 multiplies against 6), with up to 1.5x more for an out-of-order chip; a hardware engineer in the PR thread estimated a chip could still do about 18 ns per iteration, comparable to a CPU (https://github.com/SChernykh/CryptonightR ; https://github.com/monero-project/monero/pull/5126). The hashrate fell from about 1 GH/s to 140 MH/s and settled at 300 to 350 MH/s; no CN-R chip is documented in its nine months. RandomX followed on 30 November 2019.
**The mapping to v6.** CryptoNight is class B (SRAM-scale memory: layer 2 closes it, the 2 GiB dataset is 1,000 mm^2 of SRAM even at N5) and class D (the forks: v7 and v8 kept the machine's shape and the chips returned; CN-R changed what the machine had to be, a per-block random program, and no chip came in nine months, which is the per-epoch random program Igneum ships). The number: 40x per joule for the X3 against a GPU; the Igneum analogue, the f = 0 recompute chip that holds the 256 MiB cache on die, reads 0.92x at the op budget and 1.86x per joule (`chip-model-v3.md` 5.4), and CN-R's 2.5x latency claim is the fixed-shape mixer's 3x factor in the other direction. The honest residue: Monero's chips were found by nonce pattern four months after a fork at 85 percent of the hashrate; Igneum has no detector yet.
### 2.12 X16R (Ravencoin, January 2018): a drawn order over fixed functions
**The mechanism and the chips.** Sixteen hash functions in an order set by the previous block hash, a per-block automatic change with no fork. The sequencing changed; the sixteen primitives did not, so one large FPGA holding all sixteen cores needs only per-block routing: BittWare and SQRL's CVP-13 (Xilinx VU13P) was announced on 7 September 2018 naming "X17r, X16r and TimeTravel10" (https://www.cryptoninjas.net/2018/09/07/squirrels-research-labs-and-bittware-launching-new-fpga-crypto-mining-hardware/ ; https://www.bittware.com/cvp-13 ; read 8 October 2026). An unknown pool held 10 to 20 percent of blocks in February 2019, 30 to 40 in March, 45 by 11 July; the OW Miner OW1 (September 2019, 182 MH/s at 1,500 W, about USD 500) and SKC Turing R1 (680 MH/s at 800 W, about USD 1,500) were sold as ASICs, their insides unconfirmed; against a P102-100 at 35 MH/s and 219 W the OW1 is 0.75x and the R1 about 5x per joule (https://cryptoage.com/en/1782-asics-ow-miner-ow1-and-skc-miner-turing-r1-for-the-x16r-algorithm-exist.html ; https://miningboard.com/algorithms/X16R). X16Rv2 (1 October 2019) swapped one hash in; on 30 January 2020 Tron Black called the chips' return "evident"; KawPow followed in May 2020; Ravencoin's roadmap records "ASICs have been developed for X16R (and X16Rv2)" (https://cointelegraph.com/news/ravencoin-community-clash-over-mining-algorithm-continues ; https://github.com/RavenProject/Ravencoin/blob/master/roadmap/README.md).
**The mapping to v6.** Class D in its purest form, and the warning for layer 1: a draw over a FIXED set is a configuration to a chip that holds the set. Igneum's op-mix draw is over twelve families a chip holds from genesis; what the draw does cost a chip is nothing, and what it costs the honest card is the per-vendor energy table. Does v6 close it: the fixed-function lane is closed by the per-epoch program (there is no sixteen-core pipeline to route), and the draw itself is a governance device. The number: 5x for the one X16R box with a plausible chip inside; the GPU-without-graphics sequencer's gain on Igneum's program side is bounded by that kind of figure, and its memory side by class C's 5.1x.
### 2.13 Lyra2REv2 (Vertcoin, August 2015): FPGA first, then the Zig Z1
A memory-hard sponge (Lyra2 at T = 1, R = 8, C = 256, p = 1: a matrix small enough for on-die SRAM per core, on the order of 200 KB, approximate) inside a chain of hashes. The FPGA came first: an academic Lyra2 core on 16 July 2018 and a full FPGA miner at 2.6 to 3.7 MH/s and 323 to 432 nJ per hash, "significantly more energy efficient than both a GPU and a commercially available FPGA-based miner", which confirms commercial bitstreams before the chip (https://arxiv.org/abs/1905.08792); the Dayun Zig Z1 on 19 September 2018: 6.8 GH/s at 1,200 W, 28 nm, USD 8,000, "equivalent to 100 GeForce GTX 1080 Ti", about 20x per joule against a 1080 Ti at about 68 MH/s and 250 W (https://cryptoage.com/en/1231-first-asic-miner-lyra2rev2-dayun-zig-z1.html; read 8 October 2026). Vertcoin 0.14.0 forked at block 1,080,000 (1 February 2019) "to rid the network of the current generation of Lyra2REv2 ASICs and FPGAs" with Lyra2REv3 (R = 32, p = 4: 16x the memory), then Verthash in January 2021 (https://github.com/vertcoin-project/vertcoin-core/releases/tag/0.14.0). The 22 reorgs of October to December 2018 and the December 2019 attack came through rented hash on a hashrate the forks had reset (the 5 October file, row 7). Mapping: class B (layer 2 closes it) and lesson 5 (the fork reset the hashrate to a rentable size, which v6's layers 1 and 3 never do). The number: 20x; Igneum's f = 0 chip 1.86x.
### 2.14 Eaglesong (Nervos, November 2019) and Blake3 (Alephium, November 2021): compute, embraced
Eaglesong: a new sponge; the Toddminer C1 in February 2020 at 4 months, the Antminer K5 (April 2020, 1.13 TH/s at 1,580 W), the K7 at 63.5 TH/s and 3,080 W; about 100x per joule against an RTX 2080 Ti at about 1.5 GH/s and 220 W (approximate) (https://www.asicminervalue.com/miners/bitmain/antminer-k5-1130gh ; https://miningboard.com/algorithms/Eaglesong). Blake3: double Blake3, chosen as chip-friendly; the Goldshell AL-BOX (May 2024, 360 GH/s at 180 W), Antminer AL1 (15.6 TH/s at 3,510 W) and AL1 Pro (August 2024, 16.6 TH/s at 3,730 W), AL3 (8 TH/s at 3,200 W); about 100x to 220x per joule against an RTX 4090 at 6.0 GH/s and about 300 W (https://www.asicminervalue.com/miners/goldshell/al-box ; https://whattomine.com/asics/294-bitmain-antminer-al1-pro). Blake2s on Kadena: the Goldshell KD5 (March 2021, 18 TH/s at 2,250 W) and Antminer KA3 (166 TH/s at 3,154 W) against a GTX 1660 Super at 633 MH/s and 79 W: about 1,000x and 6,500x (https://www.asicminervalue.com/miners/goldshell/kd5 ; https://miningboard.com/algorithms/Blake%20%282s-Kadena%29). NexaPow (SHA-256 plus a secp256k1 Schnorr signature per nonce, "useful ASICs"): the DragonBall A21 (December 2024, 3.4 GH/s at 1,800 W, USD 6,999) is only about 2.2x per joule against an RTX 4090 at 320 MH/s and 380 W, because big-integer elliptic-curve arithmetic leaves a fixed pipeline little to strip (https://spec.nexa.org/mining/NexaPOW/ ; https://www.asicminervalue.com/miners/dragonball-miner/a21 ; all read 8 October 2026). Mapping: class A, closed since class v2; the numbers (100x to 6,500x) are what a fixed function costs, and NexaPow's 2.2x is the one compute-bound row where the chip's edge stayed small, because the work was already a wide multiplier chain, which is the shape of the ALU shadow's k band (0.3 to 0.8) in another chain's numbers.
## 3. The three lessons, with the evidence
**Lesson 1. The chip that stores the dataset is firmware-immune to every draw; only joules and memory growth move it.** Evidence: the five Ethash chips (section 2.1) never touched Ethash's compute and never needed to; Rao's audit said in 2019 that "the energy expended to move data from DRAM to compute are the same for GPU or ASIC" and that the only lever a chip has is the memory's own energy per bit, which Jasminer took by bonding the DRAM to the die; the Least Authority audit's conclusion that "the random math core likely prohibits the build of a light-evaluation based ASIC" was about the f = 0 chip and said nothing about the f = 1 chip, which is the one that shipped on Ethash in three forms. On Igneum's side the identity of `counter-asic-4-research.md` section 2 says it in one line: at zero premium the edge is E_card over E_mem and no hash change touches it. What v6 inherits: 5.1x per joule on GDDR7 at zero premium against the 5090 unlocked, 3.6x at its 1,300 MHz knee, 2.1x with the class v4 shadow at k = 1, and USD 2.8 against 14.7 per MH/s. What binds on v6: every per-era draw of layer 1 and every family epoch of layer 3 must be priced against this chip as a configuration change (a recompile per epoch, a new mapping per era) and never claimed as a cost to it; the public text's "under 2x" stays worded against the recompute chip, as the 6 October verdict already requires.
**Lesson 2. Automatic change beats the human fork only where it costs the chip a redesign, and the one parameter that does is the memory.** Evidence: Grin's three Cuckaroo tweaks (17 July 2019, 15 January 2020, 16 July 2020) each changed the edge function and no chip ever shipped for the lane, but each was a hard fork with a new solver and the lane was scheduled to die; Monero's v7 and v8 kept the machine's shape and the chips were back inside four months, CN-R changed the machine and no chip came in nine months, RandomX changed it again and the X5 took 46 months; X16R's per-block order cost the FPGA nothing; Scrypt-N's public, slow growth was abandoned before a chip because the chip could be sized for years of it; Ethash's DAG growth is the one automatic rule in the record that killed a shipped chip, and it killed the one chip whose memory was sized to the card fleet's own limit (4 GB), 20 to 27 months after shipping. The compile and design cycles bound the race: Bitmain built the A3 in about 5 months and Halong the B52 in 9 (Vorick), a full Vivado compile on a mid-size part runs 42 to 160 minutes and hours on a large one (PRflow, FPT 2019), so an hourly program outruns every compile and a six-monthly change outruns no chip. What binds on v6: layers 1 and 3 close the governance failure (no fork, no hashrate reset to a rentable size, which cost Vertcoin two 51 percent attacks) and tax a sequencer chip die area, not architecture: Rao's own figure is about 1 M gates and 0.025 mm^2 at 10 nm for ProgPoW's whole inner loop, so a lane array carrying every reserve family is the class v4 shadow core's 30 mm^2 of N5 and USD 25 to 40, not a wall. Layer 2 is the one real lever, and its value is its floor and ceiling, not its tracking: the floor keeps class B closed (2 GiB is 1,000 mm^2 of SRAM at N5), and a ceiling under the honest tiers' memory is a requirement, because any rate that ages out a 32 GB chip board retires the 8 GB card first.
**Lesson 3. A steered or biased address pattern is always found after launch unless the test lives in the acceptance rule, and every new draw needs its own null.** Evidence: Kik's exploit came five months after two audits that named the seed's keccak as a thing to scrutinise and three days after the EIP was declared dead; Dinur and Nadler found MTP's under-1 MB proof before launch only because the construction was published and reviewed, and the fix was a construction change; AP-F8-1 was found by the attack-pass lane one day after class v4 reached the devnet, in 96.6 percent of the class's programs, and took three sub-versions and class v5's floor to close; the in-house pass then found a program that passed every part of the rule and still read a live hot set, which is why the floor sits at 0.995 and not 0.98. What binds on v6: layer 4's test is a per-site ratio against the window model, and the window model is a function of the dataset size, the windows, the era stride and the read width, so a draw of any of those changes the null and the census must be re-derived per era (2.2 s per candidate at 2^20 on one box core; 0.7 percent of candidates refused at the floor); the residue the floor cannot reach without refusing most clean programs (the shadow-written concentrations at 0.9992 to 0.9997, about 1.0004x) moves with any draw of the shadow placement and needs its own ceiling per era, with the F8 gate's 1.2x-of-window shape.
## 4. The four layers against the history, layer by layer
### 4.1 Layer 1: per-era draws of the class parameters
What is declared: the parameters now fixed by release (the mixer round count within the tested margin, the op-mix weights within the measured safe band, the read width, the program length, the shadow placement) are drawn per era from chain state like the program. What the history says about each:
| Parameter drawn | The precedent | What the draw costs a chip | What it costs the honest card and the verifier | Reading |
|---|---|---|---|---|
| Mixer round count (within the tested margin) | RandomX made the item derivation itself a random program so a chip could not hard-wire it (SuperscalarHash); CryptoNight-R's random math raised chip latency 2.5x | nothing on the f = 1 chip (it derives no item); on the f = 0 recompute chip the fixed shape is the 3x factor, and a drawn ROUND COUNT keeps the shape: the chip builds the widest count and gates the rest | the verifier's 10 ms gate caps the count (x8 is 2.1 ms per warp on the reference core, x16 about 3.7); the daily build 23 to 77 ms at x8 | a draw of the count within a margin the chip already covers is firmware; the lever against the recompute chip is a drawn SHAPE (the 5 October addition 2, reserve), and the recompute chip is not the one anyone builds (chip-model-v3 5.6) |
| Op-mix weights (within the measured safe band) | X16R drew the ORDER of sixteen fixed hashes per block and an FPGA served it at 1.3x within 20 months; ProgPoW drew the math per period and no chip exists on its adopters in eight years, on small prizes | a sequencer chip over the twelve families covers any weight table; the measured GPU cost per family is the real constraint (shfl 55.8 pJ against add 11.3 on the 5090: a shuffle-heavy draw taxes the card up to 5x per instruction with no better k) | the per-program hash-rate spread must stay under the 5 percent rule on every vendor across the band (the six-era spread was 1.3 to 3.2 percent on the era layout) | right as a governance device (no fork), neutral as a chip device; the band must be bounded by the per-vendor cost table of `counter-asic-4-research.md` 15.1a, not only by the rate spread |
| Read width | w16 moved the honest denominator 2.7 percent and the chip's cost not at all; w64 made the 5090 bandwidth-bound (71.9 MH/s); the 9070 XT pays a 64-byte line at every width | the chip pays the same 32-byte atom at w4 and w16; wider reads hand a custom controller the Ren-Devadas bandwidth lever (the Ethash chips' whole edge) | a 47 percent loss on the 5090 at w64 | the one parameter whose draw can move the memory physics the wrong way: keep the allowed set at {1} (as 1.13.1 already does) unless a wider width is measured latency-bound on all three vendors; a draw over {4 B, 16 B} is harmless and worthless |
| Program length | ProgPoW's loop count and Ethash's 64 accesses were fixed; RandomX's 8 chained programs of 256 instructions fixed; no chain drew its program length | a longer program is more shadow work per hash: the class v4 lever (the latency ladder) priced at 2.1x at k = 1; a chip builds the longest rung's core and idles it on short eras | the verifier rung 3 is inadmissible on the reference core (latency-ladder section 5); the card's premium per op 6.2 to 11.3 pJ measured | the draw must stay inside the admissible rungs (0 to 2); its value is the governance one (the ladder stepped by draw instead of by 90 percent signal), and the honest card pays the premium on every era |
| Shadow placement | no precedent in any chain; the per-load placement (16 blocks of 16 after every load) was DEAD as drawn on 7 October (acceptance in execution order accepts 1.4 percent of candidates; `counter-asic-4-research.md` 20.2a-close) | the per-load form would force the chip's ALU core inside every read's dependency (the USD 200 M break-even row, 16.2) if a sound form existed | compile-ahead at 16 sites; a class change | draw only over placements shown sound (today: the one block after the loads); the per-load form is the research item, not a draw value |
### 4.2 Layer 2: the dataset's size tracking chain-state growth with a floor
What is declared: the state-derived dataset's size tracks chain-state growth, with a floor, so fixed-memory silicon ages out. The history has exactly two scheduled memory-growth rules that ran against shipped hardware, and one of them killed a chip:
| Precedent | The rule | What it did to chips | What it did to honest cards | Source |
|---|---|---|---|---|
| Ethash DAG growth | +8 MB per epoch of 30,000 blocks (about 0.7 GB a year; Rao's audit, slide "DAG size"); 1 GB at launch (July 2015), 3.0 GB by July 2019, 3.94 GB by November 2020 | the Antminer E3 (shipped July 2018, 4 GB of DDR3) ran out of DAG room on Classic at epoch 328 (about 3.56 GB, March 2020) and on Ethereum at about block 11.4 million (about October 2020) after a 30 March 2020 firmware stretched its DDR use: 20 to 27 months after shipping, 57 to 63 months after the hash went live; the A10 Pro (6 GB) and every later chip carried more memory than any card of its year and never aged out | the same rule retired 3 GB cards in 2018 and 4 GB cards by December 2020; Ethereum Classic cut the DAG to 2.47 GB (Thanos, ECIP-1099, block 11,700,000, 28 November 2020) to keep the 4 GB cards after the August 2020 51 percent attacks, and its own text says the fork kept "competitive mining across both GPU and ASIC hardware" | https://ethereumclassic.org/blog/2020-11-27-thanos-hard-fork-upgrade/ ; https://coingeek.com/memory-limitations-prompt-bitmain-antminer-e3-to-halt-etc-support/ ; https://cointelegraph.com/news/bitmains-antminer-e3-to-continue-mining-ether-with-new-update (all read 8 October 2026); Rao's audit, the DAG-size slide (https://github.com/ethcatherders/progpow-audit) |
| Autolykos v2 table growth | N = 2^26 elements of 31 bytes (2.08 GB) until block 614,400; then about 5 percent every 51,200 blocks (N doubles every 102,400 blocks, about 142 days at Ergo's 2-minute block) to a cap of 2,143,944,600 elements at block 4,198,400 (about 66 GB) | no chip has shipped for Autolykos as of today, on a small prize; the rule has never been tested against one | the table passes 8 GB in about 2 years of growth and 16 GB a year later (arithmetic on the rule); Ergo's GPU fleet thins by card memory on a schedule it chose | https://docs.ergoplatform.com/mining/autolykos/ (read 8 October 2026) |
The arithmetic that binds. A chip's memory is bought by the board, and today's prices are the chip model's: GDDR7 about USD 20 per 2 GB device, so 32 GB on a 512-bit board is USD 320; one HBM3 stack 24 GB about USD 200 (`chip-model-v3.md` 5.1, September to October 2026 prices). Igneum's schedule as specified (2 GiB plus 0.5 GiB a year, doubling steps at years 4, 12, 28; spec 1.13.3) reaches 4 GiB at year 4 and 8 GiB at year 12. Against that schedule:
| Memory the chip or card holds | Years until the dataset passes it | Who it is |
|---|---|---|
| 8 GB | 12 | the honest 8 GB card (the first tier out, `card-lifetime-2026-10-05.md`) |
| 12 GB | 20 | the honest 12 GB card |
| 16 GB | 28 | the 9070 XT class |
| 24 GB | 44 | one HBM3 stack; the 4090 and the M5 Max class |
| 32 GB | 60 | the f = 1 GDDR7 chip's board; the 5090 |
So layer 2 as a rate "tracking chain state" ages out fixed-memory silicon only if the dataset grows faster than a chip generation's memory headroom, and every rate that does that retires the honest small cards first by the same table. The E3 is the only case in the record where a growth rule beat a chip, and it beat a chip that had under-provisioned memory by a factor the card fleet also hit (4 GB). The rule that would hurt the f = 1 chip is one that keeps the dataset above what one board of commodity DRAM holds at the chip's price point, and that rule is unaffordable for the honest fleet. The honest reading: layer 2 is the right lever class (the memory is the one parameter a stored-dataset chip cannot read as firmware), and its value is set by the floor and the ceiling, not by the tracking: a floor keeps the dataset above every SRAM die (class B stays closed: 2 GiB is 1,000 mm^2 of SRAM even at N5), and a ceiling keeps it under the honest tiers' memory. Between those two lines the chip's board holds whatever the card holds, and the growth rate changes nothing for it. What "tracking chain state" adds over the fixed schedule is governance (no release decides the size) and the class v5 link (the dataset is built from the state, so the size follows the state's record count naturally); it is not an anti-chip rate. Resolved by the synthesis lane (11:2x UK, `docs/design/class-v6-rotating-family.md` section 7b on `counter-asic-4`): layer 2's rule carries a per-year ceiling as its first constant, the fixed schedule's power-of-two step for that year, so the chain's state can bring a step forward and never add one; "retires cards before chips" is recorded as the reason the ceiling exists. Under that rule the table above is the ceiling's own schedule, and the honest 8 GB tier's year-12 line stands whatever the state does. What this file adds for the record: if the chain's state grew the way Ethereum's did (approximate, from memory: the account and storage state passed 100 GB in its eighth year), a dataset tracking it with no ceiling would have outgrown every consumer card inside the first decade; the ceiling is what makes layer 2 a governance rule and not a fleet-retirement rule.
### 4.2a One schedule, priced (main's ask of 11:2x UK, for the 20:00 reading)
Main asked for one dataset-floor schedule priced per tier and against the chips, so the 20:00 reading carries a decision: 6 GiB at the v6 epoch, 10 GiB two years on, 14 GiB at four years, each step by height like a class epoch, the schedule a consensus field. The method is `card-lifetime-2026-10-05.md`'s: a card's room for the dataset is its usable memory (75 percent of card memory, the standing rule "the whole working set under 6 GB on an 8 GB card"; 50 percent of Apple unified memory) less the non-dataset working set (best: 32 KiB scratch, the cache freed after the build, 254 to 479 MiB by tier; worst: 128 KiB scratch, the cache resident, 600 to 1,500 MiB). The population is the bench table's measured cards (`site/miner-bench.json`, 8 October 2026: 32 distinct consumer cards, one Apple part, eight datacentre parts); the fleet's own card census is not a file this lane could find and is owed, so the shares below are by count of distinct measured cards, not by hashrate.
**What the memory arithmetic says about main's three steps.** A step fits a tier when the dataset plus the working set is under the tier's usable memory.
| Step | 8 GB (usable 6,144 MiB) | 10 GB (7,680) | 11 GB (8,448) | 12 GB (9,216) | 16 GB (12,288) | Apple 16 GB unified (8,192) | 24 GB (18,432) | 32 GB (24,576) |
|---|---|---|---|---|---|---|---|---|
| 6 GiB (6,144 MiB) | **does not fit the rule**: 6,398 best, 6,744 worst, 78 to 82 percent of the card; fits only if the budget rule moves to about 85 percent (a headless rig with no display) | fits (6,404 best, 63 percent) | fits | fits | fits | fits (6,432 MiB, 79 percent of its usable half) | fits | fits |
| 10 GiB (10,240 MiB) | out | out | out | **out**: 10,506 best, 10,888 worst, 86 to 89 percent | fits (10,518, 64 percent) | out | fits | fits |
| 14 GiB (14,336 MiB) | out | out | out | out | **out**: 14,614 best, 15,032 worst, 89 to 92 percent | out | fits (14,752, 60 percent) | fits |
So main's schedule as stated retires, under the standing budget rule: at 6 GiB the 6 GB tier and, unless the rule is loosened to about 85 percent, the 8 GB tier (22 percent of the measured consumer cards: the RTX 3070, 3070 Ti, 3060 Ti, 4060, 4060 Ti 8 GB, 5060, 2070 Super) on day one; at 10 GiB the 10, 11 and 12 GB tiers and Apple 16 GB (another 31 percent: the RTX 3080, 1080 Ti, 2080 Ti, 4070, 4070 Ti, 4070 Super, 5070, 3080 Ti, 3060, Arc B580) at year 2; at 14 GiB the 16 GB tier (28 percent: the RTX 5060 Ti, 5070 Ti, 5080, 4080, 4080 Super, 4070 Ti Super, 4060 Ti 16 GB, RX 9070 XT, RTX A4000) at year 4, leaving 24 GB and above (16 percent of the measured consumer cards, plus every datacentre part). That is not the tier order main's note describes, so the schedule that drops the tiers in that order is priced beside it:
| Step | Dataset | Year | Who falls off (share of the 32 measured consumer cards) | Who holds, and at what share of card memory (best / worst) |
|---|---|---|---|---|
| v6 epoch | 5.5 GiB (5,632 MiB) | 0 | the 6 GB tier (RTX 2060: 3 percent) | 8 GB at 72 / 76 percent (the worst case one point over the rule); 10 GB 77 percent; Apple 16 GB unified at 72 percent; everything larger |
| +2 years | 8 GiB (8,192 MiB) | 2 | the 8 GB tier (22 percent), the 10 GB RTX 3080 (3 percent: 8,452 of 7,680), Apple 16 GB unified (8,480 of 8,192) | 11 GB at 100 percent of usable (the 1080 Ti and 2080 Ti: out in practice, 6 percent); 12 GB at 69 / 72 percent; 16 GB; 24 GB; 32 GB; Apple 32 GB at 52 percent of its usable half |
| +4 years | 11 GiB (11,264 MiB) | 4 | the 12 GB tier (22 percent) | 16 GB at 70 / 73 percent; 24 GB at 48 percent; 32 GB; Apple 32 GB at 71 percent of its usable half |
| the fixed schedule's own steps beyond | 16 GiB | the year the state brings it forward, else year 28 | the 16 GB tier (28 percent); Apple 32 GB | 24 GB at 68 percent; 32 GB |
Reading, per tier, in the form main asked for (the year each falls off, and the share of today's measured cards that is): on the priced schedule the 6 GB tier falls at the v6 epoch (3 percent), the 8 GB tier and the 10 GB RTX 3080 and the Apple 16 GB laptop at year 2 (25 percent of the consumer cards plus the base Apple laptop), the 11 GB Pascal and Turing parts in practice at year 2 (6 percent), the 12 GB tier at year 4 (22 percent), the 16 GB tier at the 16 GiB step (28 percent), and 24 GB and above hold through every step priced. On main's schedule as stated the same tiers fall two years earlier each, and the 8 GB tier falls on day one unless the 75 percent rule moves. Either way this is the fleet-retirement rule the ceiling exists to bound (section 4.2): a step every two years retires a quarter of today's measured consumer cards each time.
**Against the chips, with the number.** Three chips, by what the step does to each:
| Chip | What a step costs it | The number |
|---|---|---|
| The hybrid-bonded DRAM chip sized at launch (the Jasminer X4's shape: DRAM dies bonded to the logic at tapeout, 5 GB per unit in 2021; the E3's shape in commodity form, 4 GB of DDR3 soldered to the board) | its memory is fixed at tapeout, so it dies at the first step it cannot carry, as the E3 did 20 to 27 months after shipping. On main's schedule a chip sized to the 6 GiB floor with 8 GB of bonded DRAM dies at the 10 GiB step (year 2); on the priced schedule a chip sized to 5.5 GiB with 8 GB dies at the 8 GiB step (year 2); a chip sized with 16 GB holds through year 4 on both and dies at the 16 GiB step. But the schedule is a consensus field read at genesis, so a maker sizes to the step it wants to survive: the X4's 5 GB was Ethash's DAG plus a year, chosen off a public schedule; the E3 died because its memory was sized to the CARD FLEET's limit, not to the schedule | kills only the chip that under-sizes; sizing to 16 GB costs the X4 shape about 3x its 2021 memory die area (approximate) and the E3 shape USD 160 more of GDDR7 (8 x USD 20, chip-model-v3 5.1) on a USD 470 part |
| The f = 1 GDDR7 chip on a board (the record's class C; 16 x 2 GB devices, 32 GB) | nothing until the dataset passes 32 GB: every step priced here fits its board; a bigger step buys more devices at USD 20 per 2 GB | USD 0 per step through 16 GiB; its 5.1x per joule and USD 2.8 per MH/s stand at every step |
| The SRAM store (the f = 0 recompute chip holding the cache on die) | the dataset's size costs it nothing (it derives every item); the CACHE's doubling under option C costs it capex only: 128 mm^2 and USD 46 per good die at 256 MiB, 255 mm^2 and USD 111 at 512 MiB (year 4 on the fixed schedule, or the year the state brings the doubling forward), 510 mm^2 and USD 306 at 1 GiB; and the mirror's size is what forces it onto a 7 nm or better node (the 5 October file, section 2.5) | USD 46 to 111 per die at the year-4 doubling; its edge stays 0.92x at the op budget and 1.86x per joule (`chip-model-v3.md` 5.4), unmoved by the dataset's steps |
The honest sentence for the 20:00 reading: the dataset-floor schedule is a fleet-retirement rule with a chip tax of USD 0 to 160 per unit on the chips that matter (classes C and B), and it kills only a chip whose maker ignores a public consensus field; the one precedent of a growth rule killing a chip (the E3) is a precedent of a maker sizing to the fleet, not to the schedule. If the schedule is adopted, the 75 percent budget rule and the 6 GiB floor cannot both stand for the 8 GB tier; 5.5 GiB keeps the 8 GB card inside the rule in the best case and one point over in the worst, and the headless-rig reading (about 85 percent) is what makes 6 GiB fit. The fleet's hashrate-weighted census is owed before the share column is read as a hashrate share.
### 4.3 Layer 3: scheduled family epochs by height, every 180 days, no release
What is declared: a new instruction family goes live on a height schedule, every 180 days by default, with no release (the reserve of spec 1.13.2, ordered at genesis). The record of scheduled change against chips:
| Chain | The change and its cadence | Human release needed | What it cost the chip | Outcome | Source |
|---|---|---|---|---|---|
| Grin, Cuckaroo lane | a new tweak of the edge function every six months (Cuckarood July 2019, Cuckaroom January 2020, Cuckarooz July 2020), each a hard fork; the lane's reward share scheduled from 90 percent to zero by January 2021 | yes, every time | a new edge function per tweak: a redesign, not a configuration | no chip ever shipped for the tweaked lane; the chip lane (Cuckatoo31+) got the iPollo G1 at about 4x in 23 months | the 5 October file rows 18 and 19, [S61] to [S66] |
| Monero | four algorithm forks in 20 months (v7 April 2018, v8 October 2018, CN-R March 2019, RandomX November 2019) | yes, every time | v7 and v8: a re-spin (85 percent of the hashrate vanished at v7 and chips were back at 85 percent four months after v8); CN-R: random math per block, chip latency up 2.5x; RandomX: a new class of machine, parity hardware after 46 months | the forks were events; the chips survived the ones that kept the shape and died on the one that changed the machine | rows 16 and 17; Vorick on a fork-surviving chip, [S47] |
| Ravencoin X16R | the ORDER of sixteen fixed hashes drawn per block from the previous block hash; no fork needed | no | nothing: a sixteen-core sequencer reads the order as a configuration | an FPGA at 1.3x within 20 months; the X16Rv2 fork (one hash swapped) was answered by bitstreams within weeks | row 9 |
| Ethereum ProgPoW | the random math re-drawn every PROGPOW_PERIOD (50 blocks in 0.9.2, 10 blocks in 0.9.3, about 2 minutes); the op set fixed (add, mul, mulhi, min, rotl, rotr, and, or, xor, clz, popcount) | no | a sequencer over the eleven ops and a 32-register file; the audit priced the conventional compute chip at little gain and the on-die cache chip at "<< 0.1x" energy per hash | never deployed on Ethereum; no chip on KawPow or FiroPoW in five to six years on small prizes | https://eips.ethereum.org/EIPS/eip-1057 and https://github.com/ifdefelse/ProgPOW (read 8 October 2026) |
| Igneum class v2 to v6 | a new program every epoch; era draws every 180 days; one reserve family per era | no | a recompile per epoch; the reserve families are in the shipped generator from genesis, so a chip that reads the reserve at genesis carries every family's datapath from day one | the governance failure is closed; the chip is not | this file |
The clocks that bound the race (all read 8 October 2026): Bitmain built the A3 "in about 5 months" and Halong the B52 "in about 9" (Vorick, through https://www.nextbigfuture.com/2018/05/obelisk-explains-the-state-of-asics-and-crypto-mining.html); KnC taped out a 20 nm part "only 3 months after starting the project" in 2014 (https://www.design-reuse.com/news/34090/20nm-asic-for-bitcoin-mining.html); ASICMiner went from founding in July 2012 to 64-chip boards on 31 January 2013 (Taylor, IEEE Computer 2017, https://michaeltaylor.org/papers/Taylor_Bitcoin_IEEE_Computer_2017.pdf); Linzhi from founding (February 2018) to tested boards (December 2020) took 27 months; Rao's 10 nm-class project "1+ year". A full Vivado compile on a mid-size Xilinx part (ZCU102) runs 42 minutes typical and 160 worst (PRflow, FPT 2019, https://ic.ese.upenn.edu/abstracts/prflow_fpt2019.html) and "several hours" on a large device (a 2021 Paderborn talk); the Least Authority audit called a 2-minute period "impractical" for a bitstream; the one KawPow FPGA on record is an academic VU35P build at 5.6 MH/s (NTU 2022, https://tdr.lib.ntu.edu.tw/handle/123456789/84103?locale=en), under an eighth of a 2022 GPU. So an hourly program outruns every compile and every design cycle; a 180-day family epoch outruns no chip's design cycle and does not need to, because the family is in the generator from genesis.
Reading. Scheduled change beat chips in exactly one shape: Grin's, where each change was a new function nobody could know in advance, and even then only because the lane was built to die. Every change a chip could read at genesis (X16R's order, ProgPoW's period, Igneum's reserve) became firmware. Layer 3 as declared is X16R's and ProgPoW's shape, automated and spaced at 180 days: it removes the fork (lesson 5 of the 5 October file, and the one thing that cost Vertcoin two 51 percent attacks), and it taxes a sequencer chip die area for families not yet live. The honest number for that tax is small: the reserve's candidates are integer ALU operations (shifts, bit-field extract, andn, byte permute, popcount, select, the second shuffle form; mm8 last), each a few thousand gates per lane; against the f = 1 chip's USD 470 of memory and board, a lane array carrying every reserve family is the same 30 mm^2 of N5 the class v4 shadow already forces (`counter-asic-4-research.md` 16.1: USD 25 to 40 of die). The one family that is not cheap for a chip to carry idle is one whose unit is large (mm8's tile engine), and that is exactly the family the measured rows say not to use for forcing (the 5090's int8 MAC at 1.4 to 4.1 pJ against a 5 nm array's claimed 0.04 to 0.4: `counter-asic-4-research.md` 15.1a). So layer 3 is right for governance and neutral for the chip; what would make a family epoch cost a chip a redesign is a family whose semantics are not knowable at genesis, which is the random item-derivation program of the 5 October addition 2 (a per-day SuperscalarHash-style derivation), the one RandomX idea Igneum has not taken, and it acts on the f = 0 chip only.
### 4.4 Layer 4: the (c''') floor and the F8-form uniformity test per era draw
What is declared: the (c''') acceptance floor (the per-site distinct-index ratio at or above 0.995 against the window model, spec 1.4.7.2) and the F8-form uniformity test generalised to each era's parameter draw, with a redraw on failure. The history has three attacks of this one class, and in every case the test that would have caught it lived outside the acceptance rule:
| Attack | The steer | Found when | What the fix was | Source |
|---|---|---|---|---|
| Kik's ProgPoW exploit (4 March 2020) | the 64-bit seed between the two keccak passes: fix a seed, compute its mix once, then grind nonces until keccak_progpow_64(header, nonce) equals the seed; the memory path is run once per 2^64 nonces, so a chip that never touches the DAG wins once the difficulty passes 2^50 | after two tentative approvals and five months after both audits (September 2019) | ProgPoW 0.9.4 widened the carried state from 64 to 256 bits (the digest of the first keccak, plus the mix, plus padding) | https://github.com/kik/progpow-exploit and https://github.com/ifdefelse/ProgPOW (read 8 October 2026) |
| Dinur and Nadler on MTP (2017) | the prover controls the memory's contents, so by injecting blocks it steers Argon2d's data-dependent addresses into a set it can hold in under 1 MB at a 170x compute penalty, in place of 2 GB | before launch, by cryptanalysis, not by the design's own test | MTP 1.2 patched the construction | the 5 October file [P18] [S38] |
| AP-F8-1 on class v4 (7 October 2026) | a load site whose source register was last written by a lossy op (or, mul, mulhi) saturates to all-ones or zero at a known rate, and the era map sends the constant to one item; 96.6 percent of class v4 programs carried a lossy-sourced load; the worst seed read 29x the window model on one item | in the attack-pass lane's census, one day after the class shipped to the devnet | sub-versions 1 to 3 (the freshness fixpoint, the executed shadow block, the (c'') ratio at 0.98), then class v5's (c''') at 0.995 | `docs/plans/counter-asic-3-status.md` section 7c; spec 1.4.7.2 |
Reading. Layer 4 is the right answer to this class and the only one of the four layers that acts on a mechanism the record shows beating hashes after launch. Two things bind on it. First, the test can only see what its null models: (c''') measures distinct-index ratios against the window model, which is a function of the dataset size, the per-site windows, the era stride and the read width; a draw of the read width or the program length changes the null, so "generalised to each era's draw" means the window model is re-derived per era and the census re-run per draw, not one floor reused. The cost is known: 2.2 s per candidate at 2^20 on one box core, once per epoch draw on a node, and 0.7 percent of candidates refused by the floor (spec 1.4.7.2). Second, the residue the floor cannot reach without refusing most clean programs is the shadow-block-written concentrations at 0.9992 to 0.9997, worth about 1.0004x to a chip (spec 1.4.7.2), and a per-era draw of the shadow placement moves that residue, so the redraw rule needs its own ceiling stated (the 1.2x-of-window gate of F8 is the right shape; the number per era is the census's to set). The Kik lesson is separate and already closed: Igneum's seed is 256 bits through the VDF and the program is the epoch's, so there is no 64-bit state to grind; the header-locality search (the 5 October check 1) remains the nearest analogue and was measured by adv-accept-2 in the in-house pass.
## 5. What to add, optimise or invent, ranked (first cut; the full report re-ranks with the other lanes' findings)
| Rank | What | Why the history says so | Cost to the honest card | Where |
|---|---|---|---|---|
| 1 | **State layer 2's floor and ceiling in gigabytes per tier, before its rate.** The floor above every SRAM die (today's 2 GiB holds); the ceiling under the honest tiers' memory on a stated glide (the 8 GB tier's life is the first number) | the E3 is the only chip a growth rule ever killed and it was the chip with the fleet's own memory limit; Scrypt-N was abandoned because its growth was public and slow; Autolykos's growth is untested; a dataset that tracks chain state literally outgrows every card if the state grows the way Ethereum's did (approximate) | none at the floor; everything at the ceiling | genesis rule, with the card-lifetime table |
| 2 | **The clock and the detector**, unchanged from the 5 October ranking and still unbuilt: the per-program rate spread and nonce pattern on the observer (the method that found Monero's chips at 85 percent), plus a share-by-key-and-template instrument (what found Qubic until it randomised), plus the issuance trigger at about USD 50 K a day | every chip in the record was on its chain before it was announced (Monero 2017, Zcash's three groups, SChernykh's 2021 reading of the X5) | none | before the public testnet |
| 3 | **The read width's draw band stops where the honest card stops being latency-bound**: the synthesis lane keeps {4 B, 16 B} on the measured warrant (w16 latency-bound within 2.7 percent on the 5090 and the 9070 XT, within 1 percent on the M5 Max); w64 is out (the 5090 bandwidth-bound at 71.9 MH/s); the draw costs every chip nothing (the same 32-byte atom at either width) and is a governance value only | the Ethash chips' whole edge was the bandwidth lever Ren and Devadas name; a width that makes the card bandwidth-bound hands the chip that lever | within 2.7 percent at w16; 47 percent at w64 | the band {4 B, 16 B} in the class v6 spec; nothing wider without a per-vendor latency-bound measurement |
| 4 | **Bound the op-mix draw by the per-vendor energy table, not only by the rate spread**: the 5090 pays 55.8 pJ per shuffle against 11.3 per add, so a shuffle-heavy era taxes the honest card up to 5x per instruction for no better k | X16R's drawn order cost the chip nothing and the fleet nothing; Igneum's draw can cost the fleet watts | up to 2x the premium per instruction at the band's edge | the band's definition in the class v6 spec |
| 5 | **Order the reserve by what a sequencer cannot fold into firmware**, mm8 last; and name the random item-derivation program (the 5 October addition 2) as the one reserve item whose semantics are not knowable at genesis | the kHeavyHash chips and the 5090's own 1.4 to 4.1 pJ per int8 MAC; RandomX's SuperscalarHash is the one idea Igneum has not taken, and it acts on the f = 0 chip only | none at launch | reserve ordering, genesis |
| 6 | **Generalise layer 4 with a null per drawn parameter**: the window model re-derived per era, the census per draw, and a stated ceiling for the shadow-written residue per shadow placement | lesson 3 | 2.2 s per candidate once an epoch on a node | the class v6 acceptance rule |
| 7 | **The per-load shadow placement as a research item, not a draw value**, until a sound construction is drawn (the 16 x 27 form accepted 1.4 percent of candidates) | it is the one placement that would force the chip's ALU core inside every read's dependency (the USD 200 M break-even row) | compile-ahead at 16 sites | `counter-asic-4-research.md` 20.2a |
## 6. Consequences per user tier
| Tier | What this history means for it | What is being done |
|---|---|---|
| Home miner, one 8 GB card | Layer 2 is the one layer whose rate reaches this tier first: at 0.5 GiB a year the 8 GB card is out at year 12, and any rate fast enough to age out a 32 GB chip board is out of this card's life inside two years. No chip of class A, B or E touches this miner under v6; the class C chip (the memory system without the GPU) reaches it as it reached Ethash's 4 GB miners: by price per MH/s, 5x | the layer 2 ceiling stated per tier (section 4.2) before the rate; the issuance clock and the detector (section 5) |
| One 12 GB or 16 GB card | as the 8 GB tier with 20 and 28 years on the schedule; the 9070 XT's 10.6 microjoules per hash is 23x behind the GDDR7 chip per joule on the model, so this tier's card is the first the class C chip displaces | nothing in the hash fixes AMD's dependent-read rate; the vendor-share metric is the warning |
| One 24 or 32 GB card (5090, M5 Max) | the honest best: 1.69 microjoules at the 5090's knee against the chip's 0.47; 3.6x at zero premium, 2.1x with the shadow at k = 1; the M5 Max at 0.78 microjoules is 1.7x behind the GDDR7 chip with no shadow at all | the operating point as the shipped default (Ember); the shadow's rung |
| A rig | the Ethash precedent in full: chips at 2x to 5x per joule and 5x per dollar took the hashrate over four years, and the ASIC share stayed small only while the chips were not cheap enough at scale; the model says this chip is (USD 2.8 against 14.7 per MH/s) | the break-even cap row (USD 100 M in years 1 to 2 with the N5 shadow core) is the real wall; the clock |
| A pool user | MoneroCrusher found chips at 85 percent of Monero's hashrate by the share pattern, four months after a fork; Igneum's detector is the same method on the observer, unbuilt | the detector before the public testnet (section 5) |
| Every tier, on governance | no fork, ever, for a draw or a family: the history's clearest lesson (Vertcoin's two 51 percent attacks after forks, Monero's four, Ethereum's two-year ProgPoW fight) is the one v6's layers 1 and 3 close outright | nothing further |
## 7. Unverified and owed
- The research lanes' returns for the chip mechanism columns (section 1 and 2) are the first cut's owed content; where a lane reports "not found" the row says so and carries the 5 October file's figure.
- The Vorick post (13 May 2018) is read through secondary coverage (davidgerard.co.uk, zycrypto.com, nextbigfuture.com, a steemit copy; read 8 October 2026): the secret-Monero-ASIC claim ("since early 2017, making up 50 percent of the hashrate"), the three Zcash groups, the Equihash fork-following architecture, "about 5 months" for Bitmain's A3 and "about 9 months" for Halong's B52, the A3 under USD 10 M with USD 20 M of orders in eight minutes, the manufacturer withdrawal that cost Obelisk "north of USD 2 million". The 5 October file's "13 months for a startup" and "a chip able to survive Monero's forks at under a 5x hit" were NOT found on any fetched page and are carried here as unverified; the Monero fork-survival fact that is verified is the record itself (chips back inside four months of v8).
- The KawPow fork block and its 3-block period are approximate (the minerstat and Tron Black pages answered 403).
- The Ethereum DAG date of passing 4 GB on the main chain (about December 2020) is approximate; the Classic figures (3.94 GB at epoch 376, 27 November 2020) are cited.
- The Ethereum state-size figure in section 4.2 is from memory, approximate, and is the open question handed to the synthesis lane.
- Every per-joule ratio for a chip against a GPU is arithmetic on the cited rate and watt figures of both and is approximate by construction.
- Nothing here is a measurement; the Igneum figures are the repo's measured rows as cited, and the chip figures are the chip model's, modelled.
## 8. Sources
Every URL is cited inline at the row that uses it, with "read 8 October 2026" at the row or the section; the 5 October file's [S], [P], [E] and [L] lists are cited by their tags and not repeated. The primary documents read in full by this lane (text extracted where the page is a PDF): EIP-1057 (https://eips.ethereum.org/EIPS/eip-1057); the ifdefelse ProgPOW README (https://github.com/ifdefelse/ProgPOW); the Least Authority audit (https://leastauthority.com/static/publications/LeastAuthority-ProgPow-Algorithm-Final-Audit-Report.pdf, report version 9 September 2019); Bob Rao's hardware audit (https://github.com/ethcatherders/progpow-audit, 6 September 2019); Kik's exploit (https://github.com/kik/progpow-exploit); RandomX design.md, design_v2.md, specs.md, PR 317, release v2.0 and issue 11 (https://github.com/tevador/RandomX); Tromp's README and the Grin forum threads named in 2.4 (https://github.com/tromp/cuckoo ; https://forum.grin.mw); the Ergo Autolykos docs (https://docs.ergoplatform.com/mining/autolykos/); the Thanos post (https://ethereumclassic.org/blog/2020-11-27-thanos-hard-fork-upgrade/); the TechInsights Jasminer notes (https://www.techinsights.com/ko/node/51986 and /52149); the Fudan NDSS 2019 paper (https://www.ndss-symposium.org/wp-content/uploads/2019/02/ndss2019_09-5_Bai_paper.pdf); Percival's lookup-gap note (https://mail.tarsnap.com/scrypt/msg00092.html); Lee and Kim on Qubic (https://arxiv.org/html/2512.01437v2); PRflow (https://ic.ese.upenn.edu/abstracts/prflow_fpt2019.html). Papers of 2024 to 2026 found: Blocki and Smearsoll, "Provably memory-hard proofs of work with memory-easy verification", ePrint 2025/1456 (Omega(N^2 / log N) cumulative memory with polylog verification: https://eprint.iacr.org/2025/1456); Condrey, PoSME, arXiv 2604.15751; Yang et al., PHICOIN, arXiv 2412.17979 (a resistance claim with no algorithm in the abstract). Pages that refused every lane (403, 404, DNS): Vorick's original post on Medium and sia.tech and its archive copy; Linzhi's and ifdefelse's Medium posts; MoneroCrusher's Medium post (figures taken from criptonoticias coverage); bitmain.com's product list; support.bitmain.com's Z9 page; minerstat; Tron Black's posts; innosilicon.global; jasminer.com (an empty shell); cryptomining-blog.com; the Yole DBI report. The session's web-search budget ran out at 11:0x UK; everything after that is direct fetches of known URLs, and "not found" in this file means not found on a fetched page.

View file

@ -0,0 +1,256 @@
# Class v6 invention lane: layers beyond the four, each priced against the chip that stores the dataset
8 October 2026, first cut 11:1x to 13:xx UK (BST, the Mac's clock; the boxes print CEST, one hour ahead, and nothing below is stated in box time), branch `class-v6-invention` from the mirror's master at 17d61d52, research lane C under the Counter ASIC coordinator. The founder's word of 11:1x UK: class v6 is declared with four layers as its spine (per-era draws of the released parameters; the state-derived dataset's size tracking chain state with a floor; scheduled family epochs every 180 days with no release; the acceptance floor and the F8-form uniformity test generalised to each era's draw), and deep research opens: "see if anything can be optimised, added or invented". This file is the invention lane's answer: every candidate layer beyond the four as one paragraph, one known-failed test stated as a harness run, and one chip-model row, then the ranked list of what v6 should add. The synthesis is the research lane's `docs/design/class-v6-rotating-family.md` (branch `counter-asic-4`), which reads this file.
Status: a research document and a gate plan. No consensus code this week; nothing here touches the devnet, the testnet object, the spec or any served page. Every number carries a label: **measured** (a run on a named box or card, the log named), **modelled** (arithmetic on the chip model's cited figures, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL and date), **approximate** (from memory or a scaling). The rule every candidate is held to: **reject what costs GPUs more than it costs chips, and say so with the number; keep what raises a chip's k or capex or shortens its useful life more than it raises every GPU tier's cost.**
## 0. One page
The frame is last night's identity (`docs/analysis/counter-asic-4-research.md` section 2): against the chip anyone builds, a GPU's memory system without the GPU (chip-model-v3 section 5.5, the `f = 1` chip), the per-joule edge is `(E_card + F) / (E_mem + k F)`, with `F` the premium per hash the card pays for forced work and `k` the chip core's energy per forced op over the card's. Zero premium is zero forcing. No hash-side design reaches a chip edge of about 2x at zero premium; the premium-free floor is the card's own whole-card energy over the chip's memory energy, 3.6x on a 5090 at its knee (measured card, modelled chip), and the class v4 shadow buys 2.1x at `k = 1` for 82 to 90 W on a 5090. The four layers of v6 render a fixed-function chip useless on its release day and move nothing in the identity for the GPU-like chip. So an invented layer can do one of four things, and only four: raise `k` (nothing on measured rows reads `k` above 1, section 15.1a of the research file); raise the chip's capex or project cost (the one lever the identity does not see); shorten the chip's useful life (layers 1 and 3 already do this for fixed silicon); or cost the honest cards less at the same `F`. Every candidate below was read against those four.
What was measured today (all on igneum-build-1, the counter-asic-4 crate at 5984ffab, logs under `/srv/builds/v6-invention/`):
| Reading | Number | Label |
|---|---|---|
| The sound per-load shadow (one pass of a 256-instruction sub-block after every load, `mx8+shl4096x1`), the form 20.2a-close named and never drew | 234 of 256 seeds accepted within the 32-attempt cap, 0.927 rejection per candidate, mean accepted attempt 9.7; the iterated 16 x 27 form on the same seeds 76 of 256, 0.989 per candidate | measured (section 3.1) |
| The same at class v4's instruction count (`mx8+shl2304x3`: 16 sub-blocks of 144, three passes, 55,296 shadow instructions per hash) | 224 of 256 seeds, 0.935 per candidate | measured |
| The verifier on the sound forms, cold warp on core 40 with core 88 loaded (the ladder's method) | shl4096x1 8.28 to 8.55 ms; shl2304x3 8.44 to 8.82; class v4's shape 8.33 to 8.63 on the same core in the same minutes; all under the 10 ms gate | measured (section 3.2) |
| The per-load prototype's acceptance under a drawn era | 0 of 256 seeds on every per-load form: 5,536 of 7,862 bias rejections name index bit 26 or 27 "set in 0 or 16,384 of 16,384", the era WINDOW's fixed top bits, an instrument fault; so 20.2a-close's "1.4 percent accepted, 42 of 64 seeds exhaust" (run across drawn eras) measured the instrument, not the construction | measured; a correction for the research file (section 3.3) |
| The two packs for the card row | `mx8_shl4096x1` (id 75ca9547da21b200) and `mx8_shl2304x3` (id bbfdfc1dcdda0b46), OVERALL PASS, in the hash lane's PC 1 job after the AMD grid (about 12:45 UK) | exported; the card row owed |
**The ranked list of what v6 should add beyond the four** (section 4 has every column):
| Rank | Layer | What it does to the chip | What it costs the cards (5090 / 5070 Ti / M5 Max) | Status |
|---|---|---|---|---|
| 1 | **Layer 5: the shadow placed per load, one pass of a long sub-block** (the capex lever: the chip's core must sit inside every read's dependency, so controller, lanes and PHY share one N5-class die or an interposer) | energy edge unchanged (`k` is `k` whichever die the core sits on); project cost about USD 30 M to about 60 M, the break-even market cap about USD 100 M to about 200 M (the mission lane's model, modelled); capex per MH/s 3.0 to 4.3 USD (modelled) | 5090: the 16 x 27 form measured 13 to 14 W UNDER the whole block at the same instruction count (448 against 462 W unlocked; 283 against 296 at the 1,300 lock), rate within 1.2 percent; the 256 x 1 form's row in today's PC 1 job / no row, the 4070's 30 W premium at its tune point as the proxy / the 16 x 27 form -0.5 percent of rate measured; the 256 x 1 form's Metal footprint OWED | acceptance measured today (0.927 per candidate, P(exhaust at 256) about 4 x 10^-9); the sub-version 3 dataflow rule in execution order is the fix that brings it toward v4's 0.68 (gate plan, section 2.1) |
| 2 | **Layer 6: register-file width drawn per era (8 to 32 registers per lane)**, the link tax on layer 5 | with layer 5, the lane state crossing the controller twice per read grows from 64 B to 128 to 256 B: 2.2 to 9 TB/s of die-to-die traffic at the 5090's read rate, past any one-stack interposer, which closes the "or an interposer" branch and forces the single N5 die (modelled); alone, nothing | 0 rate on every card while latency-bound (a GPU lane holds up to 255 registers; 7,262 lanes x 256 B is 1.9 MB against the 5090's 43 MB of register file, approximate); the verifier's register-major arrays 4x (unmeasured, under 0.1 ms by the op law); no pack form today (the register count is a generator constant) | modelled; the pack form is a generator change with its own census (gate plan) |
| 3 | **Layer 7: warp-uniform data-dependent block selection** (which of B shadow sub-blocks runs next is chosen by a warp-reduced register value, uniform across the 32 lanes, so no divergence) | nothing for the GPU-like chip (a sequencer already); an FPGA overlay or a fixed pipeline must hold all B blocks for one block's throughput: B x the shadow's LUT area (approximate); shortens a per-epoch bitstream's worth | one uniform indirect branch per iteration: about 0 (the `sel` register already does this for the immediates); verifier 0 | modelled; a generator change behind a pack (gate plan); rank 3 because it moves the FPGA lane only |
| 4 | **The reserve ordered by hardware orthogonality** (layer 3 as it stands, with the order fixed: shuffle-crossbar families, then byte-permute, then popcount and priority encoder, the int8 tile last) | a chip pre-wires every family for about USD 4 of N5 (modelled); the order makes the first unlocks the ones a 12-op datapath lacks most | 0 at 4 points (measured step costs under 1 percent of ALU time on every vendor) | an ordering rule inside layer 3, not a new layer; the history's addition 6 |
Rejected with the number, each in section 2: per-lane data-dependent branches (divergence costs the card, a chip nothing); reads tied to the shard proof per block (a refresh per block is 0.3 to 1 W on a chip, 1.3 percent of a 4090's hash time); randomised memory topology (a chip's address decoder permutes its lines for nothing; the stride and interleave are already drawn); the VRAM-size ratchet as a lever (a chip buys 24 to 32 GB that the 8, 12 and 16 GB tiers cannot: it retires cards first); proof-carrying hashes sampled by the pool (the chip holds everything the witness proves); prover-gated eligibility (proving is 1.6 kW network-wide at any hash rate, 0.7 percent of the hash's energy at 100 GH/s); time-locked parameter commitments beyond the era VDF (the drawn band is firmware; the 2-hour lead already denies the fixed chip 180 days); a fraction of reads derived from the cache (the 5090 loses about 30 percent of rate, the chip 7 percent of energy); row-straddling reads (the w64 regime, 47 percent of the 5090's rate); a refresh per block (dead by arithmetic, the research file's row 7).
Per tier, in one line each: a home miner on any card sees no change from anything here today (nothing ships; the devnet pays nothing); the 5090 tier's one number is that the per-load placement costs it LESS than class v4's whole block at the same work (13 to 14 W measured on the 16 x 27 form, the 256 x 1 form's row due about 12:45 UK); the 5070 Ti has no measured row in this lane (the 4070's rows are the proxy, approximate); the Apple tier's open question is the inline footprint of a 4,096-line block per iteration (the 1,024-line block cost the M5 Max 17 percent on 6 October, measured), which decides whether rank 1 needs a block-shape cap for Apple; a pool user sees nothing; a node verifies the per-load forms in the same 8.3 to 8.8 ms as class v4 (measured); a chip maker sees its project forced onto one advanced die by rank 1 and its interposer escape closed by rank 2.
## 1. The frame, and what a candidate must do
| Term | Value, 5090 at the 1,300 MHz knee | Label | Source |
|---|---|---|---|
| `E_card`, class v3 | 1.67 microjoules (127.3 MH/s at 213.0 W; 134.6 at 223.3 on the efficiency pass) | measured | research file 20.3 and the status file's efficiency pass |
| `E_mem`, the `f = 1` GDDR7 chip | 0.466 microjoules (16 devices, 166 MH/s at 78 W); one HBM3 stack 0.321 | modelled | chip-model-v3 5.4 |
| `F`, the class v4 shadow's premium | 0.652 microjoules (82.8 W over 126.9 M hashes x 102,100 ops: 6.4 pJ per counted op) | measured | research file 20.3 |
| The edge at zero premium | 3.6x (GDDR7), 5.2x (one HBM3 stack) | measured card, modelled chip | section 2 of the research file |
| The edge with the shadow | 2.1x at `k = 1`, 2.9x at `k = 0.5`, 3.5x at `k = 0.3`; the honest `k` band for an ALU-shaped core 0.3 to 0.8 | modelled on measured rows | 20.4 |
| The chip's capex | USD 470 of memory, controller and board per 166 MH/s: 2.8 USD per MH/s; plus the shadow core USD 25 to 40 (3.0 to 3.1); plus an interposer USD 200 (4.3); a 5090 at MSRP 14.7, at the 2026 street price about 29 | modelled; the card price measured | 16.1 |
| The project cost and the break-even cap | a 28 nm controller USD 5 M (cap about 17 M); plus an N5 shadow core USD 30 M (cap about 100 M); the core forced onto the controller's die or an interposer USD 60 M (cap about 200 M) | modelled (the mission lane's model, s = 0.30) | 16.2 |
The four doors, and the one each candidate must walk through:
1. **Raise `k`.** Measured on the 5090 (15.1a): the int32 ALU op 6.2 to 11.3 pJ against a 5 nm SIMD array's 2 to 5 (k 0.3 to 0.8); the shuffle 29 to 56 pJ against a crossbar's about 20 (k 0.4 to 0.7, and the card pays 5x the add per op); the int8 tile 0.8 to 4 pJ per MAC against an array's 0.04 to 0.4 (k 0.03 to 0.3); an L2 hit 1.4 to 2.4 nJ against on-die SRAM's 0.2 to 0.5 (k 0.1 to 0.3). Nothing reads above 1. A candidate that claims to raise `k` must name the block and the measured GPU cost per op it rests on.
2. **Raise the capex or the project.** The hash's dataset is a card's worth of DRAM and its work a fraction of a card's logic, so per-unit capex cannot pass about 4.3 USD per MH/s (16.1); the project cost is the lever that moves the break-even cap, and the only mechanism found for it is forcing the shadow core into every read's dependency (16.2). A candidate here is priced by which die it forces.
3. **Shorten the useful life.** Layers 1 and 3 kill fixed silicon at the first draw outside its wired value; a GPU-like chip's life is its memory's and its node's. A candidate here must move the GPU-like chip, or it is layer 1 again.
4. **Lower the honest card's cost at the same `F`.** The operating point (the knee lock) is the miner's lever, not the protocol's; a protocol lever here is a shape that runs cheaper per instruction on the card (the 16-instruction block effect, 13 to 14 W measured) at the same chip cost.
## 2. The candidates
Each: the paragraph, the known-failed test as a harness run, the chip row, the verdict. The harness names: `v6inv-census.sh` is `/srv/builds/v6-invention-census.sh` on build-1 (a copy sits in this lane's scratch and lands under `tools/attack/v6-invention/` with the 09:00 report), the counter-asic-4 crate's `igneum-pow accept --class <form>` over seeds `igneum-v6inv/<i>`; the F8 census is `tools/attack/f8-uniform` (`attack-f8 warps`, master); the chip row is the model's arithmetic with the GPU side measured where a pack ran.
### 2.1 Layer 5: the shadow placed per load, one pass of a long sub-block (KEEP, rank 1)
The class v4 shadow runs 256 instructions 27 times after instruction 63 of every iteration, where the next iteration's 64 base instructions and 16 loads stand between the block and every load, and a chip may run it on a second die with 64 B of lane state crossing once per iteration (140 GB/s at the 5090's read rate, a PCB link; 16.2). Placed inside every read's dependency the same work makes the lane state cross twice per read, 2.2 TB/s, an interposer-class link or one N5 die carrying controller, lanes and PHY, which the mission lane's model prices at a project of about USD 60 M against 30 M and a break-even cap of about 200 M against 100 M. The 16 x 27 form (16 sub-blocks of 16, each iterated 27 times) was built on 7 October and closed the same night: 27 passes of a 16-instruction map right before a load collapses or biases the load's address register before any base instruction can re-randomise it, and the acceptance rule in execution order refused it. The sound form named there and never drawn is one pass of a long segment per load: 16 sub-blocks of 432 instructions, each run once, the same 6,912 instructions per iteration. The experimental class caps the per-load block at 4,096 instructions, so today's census takes the two forms inside the cap that bracket it: `mx8+shl4096x1` (16 x 256, one pass, 59 percent of class v4's work per hash) and `mx8+shl2304x3` (16 x 144, three passes, class v4's exact count), and the constant-work ladder between them (16 x 256 x 1, 32 x 128 x 2 ... 16 x 16 x 16) that reads how the acceptance rate depends on the sub-block length against the pass count. The 432 x 1 form itself needs the cap raised, a one-line change in a research-only parser, which this lane does not make this week; its number is bracketed by the 256 x 1 and 144 x 3 rows and the ladder's flatness between them.
Known-failed test, as a harness run: `ERA=none SEEDS=256 /srv/builds/v6-invention-census.sh` on build-1 under `lease pool 48`. Fail: a form whose per-candidate rejection is 0.98 or above (P(exhaust) at the 256-attempt cap above 0.5 percent, an epoch without a program every few months). The 16 x 27 form reads 0.989 and fails (section 3.1). Pass: a form under 0.95 (P(exhaust) under 2 x 10^-6). The 256 x 1 form reads 0.927 and passes; so does every form with a sub-block of 36 instructions or longer (0.927 to 0.949). The second known-failed case is the instrument's own: `ERA=era` on the same seeds reads 0 of 256 accepted on every per-load form, with the window's fixed top bits named as biased index bits (section 3.3); the fixed instrument (bits at or above `28 - k_off_s` excluded from the value-level test) must accept per-load forms under drawn eras at the no-era rate within the binomial band, and still refuse the per-load record of 7 October (candidate 0 of the first export, 1,482 duplicate lanes).
Chip row:
| Column | Value | Label |
|---|---|---|
| A fixed-function chip's `k` | unchanged (the work is the same ALU mix); its capex: the controller cannot be a 28 nm part with the core elsewhere, so the project moves from USD 5 M (no core) or 30 M (a core on its own die) to about 60 M (one N5-class die or a 2.5D package); break-even cap about USD 200 M in years 1 to 2 (s = 0.30) against about 100 M | modelled (16.2, the mission lane's N3 single-die row; a GDDR7 PHY on N5 is unpriced) |
| A GPU-like chip's per-joule edge | unchanged: 2.1x at `k = 1`, 3.5x at `k = 0.3` at the knee; its capex per MH/s 3.0 to 4.3 USD against 2.8 | modelled |
| RTX 5090 | the 16 x 27 v2 export: 135.90 MH/s at 448.3 W unlocked against the whole block's 137.51 at 462.2 (the premium 137 against 151 W, -14 W); at the 1,300 lock 126.04 at 282.9 against 126.93 at 295.8 (-13 W); the 256 x 1 and 144 x 3 forms in today's PC 1 job (the hash lane, after the AMD grid, about 12:45 UK) | measured (research file 20.3); the sound forms' rows owed |
| RTX 5070 Ti | no row (no card in this lane); the 4070's class v4 premium at its tune point, 30 W for no rate, less the block effect, is the proxy | approximate |
| Apple M5 Max | the 16 x 27 v2 export 26.88 MH/s against 27.01 (-0.5 percent, Metal packbench, 7 October); the 256 x 1 form's inline text is 4,096 shadow lines per iteration where the 1,024-line block cost the M5 Max 17 percent (6 October, measured), so its footprint is the open Apple number: OWED (a Mac measurement under the measure lock, which this lane does not run; the hash lane's or the shipper's Metal row) | measured for 16 x 27; the sound form's row owed |
| The verifier | 8.28 to 8.55 ms cold with the sibling loaded (256 x 1), 8.44 to 8.82 (144 x 3), against class v4's 8.33 to 8.63 on the same core in the same minutes; the acceptance's dynamic test 35 ms per candidate (1,111 ms for 32), 13.7 candidates per seed on average: about 0.5 s of one core per epoch | measured (section 3.2) |
| The acceptance | 0.927 per candidate (256 x 1), P(256 consecutive rejections) 0.927^256 about 4 x 10^-9 per seed; class v4 sub-version 3 reads 0.681 and 2 x 10^-43; the gap is the missing dataflow rule (the per-load class is not the class v4 shape, so (a'), (c') and (c'') do not run on it; the value-level bias test catches the same population: 1,563 of 2,329 bias rejections name index bit 0 at a one-count near 4,096 or 12,288 of 16,384, the product's low-bit law) | measured (section 3.1) |
Verdict: KEEP as layer 5, the first thing v6 adds beyond the four, because it is the only mechanism found that moves the project cost, it costs the 5090 less than class v4's own block (measured on the 16 x 27 form; the sound form's row today), and its acceptance is now a measured 0.927 with a named fix (the sub-version 3 dataflow fixpoint run over the real execution order, base and sub-blocks interleaved, as the generator's draw rule) that the research file's 20.2 already asked for. What it does not do: move the energy identity by one joule. The founder's "useless as soon as it dropped" is layer 1's and 3's sentence; layer 5's sentence is "the chip that can be built costs twice as much to start".
### 2.2 Layer 6: the register-file width drawn per era, the link tax on layer 5 (KEEP, rank 2)
The hash runs on 8 registers per lane, a prototype value to be fixed at gate 1 (spec 1.4). The lane state a chip must carry is those 8 words plus the nonce and counter, 64 B, which is why the chip's 1,172 lanes are 73 KB of SRAM and why, under layer 5, the per-read crossing is 128 B at 17.5 G reads per second, 2.2 TB/s: an interposer carries that (a one-stack HBM package moves about 0.8 to 1.2 TB/s of memory traffic and a die-to-die link of a few TB/s is a 2.5D product; approximate, from memory), so the chip has an escape at USD 200 of package instead of one die. Draw the register count per era from {8, 16, 32} (the acceptance rule's part (b) over every register; the program length scaled so that every register is written, or registers above 8 initialised and read by the shadow alone) and the crossing is 128 to 256 B per lane per read: 4.5 to 9 TB/s, past the interposer class, so the single die is forced and the project's USD 60 M row has no cheaper branch. A GPU pays nothing in rate while latency-bound: a CUDA lane holds up to 255 registers, the 5090's 7,262 lanes in flight at 256 B are 1.9 MB against about 43 MB of register file (170 SMs x 256 KB; approximate), and occupancy at 32 live registers plus the kernel's temporaries fits the 64K-register SM at full residency (approximate, unmeasured). The verifier's register-major arrays grow 4x (4 KB per warp) and the interpreter's cost per op does not move.
Known-failed test, as a harness run: the generator with `REGISTERS` as a class field (a research-only change behind a pack name, `mx8+r32`), then `v6inv-census.sh` over the per-load forms at 8, 16 and 32 registers. Fail: the 32-register form's per-candidate rejection above the 8-register form's by more than the binomial band (more registers, more cold registers, more (b) rejections unless the program length scales). Pass: rejection at or under the 8-register form's; the F8 census at 2^24 on 64 seeds within 1.2x of the window model (the top 0.1 percent of items); the cold verify on core 40 with core 88 loaded under 10 ms. Then the card: one pack per register count on the 5090, rate within 1 percent of the 8-register pack at both states (the occupancy claim measured, not argued).
Chip row:
| Column | Value | Label |
|---|---|---|
| A fixed-function chip | with layer 5: the lane state per read 128 to 256 B, 4.5 to 9 TB/s of die-to-die traffic at the 5090's read rate; the interposer branch (USD 200 of package) closed, the single N5-class die forced; the project about USD 60 M either way, but with no cheaper escape; alone (without layer 5): nothing, the state crosses once per iteration | modelled, the link figures approximate |
| A GPU-like chip's per-joule edge | unchanged; its lane SRAM 73 KB to 300 KB (nothing) | modelled |
| RTX 5090 / 5070 Ti / M5 Max | 0 rate while latency-bound (approximate: the occupancy arithmetic above; a measurement is the pack); watts: the same `F` (the same ops) | approximate until the pack runs |
| The verifier | 4x the register arrays per warp (4 KB); cost per op unchanged (0.1 ns per lane-instruction, the shadow's law) | modelled |
| The acceptance | part (b) over 16 or 32 registers needs the base program to write every register: at 64 instructions over 32 registers about 13 percent of registers are never written (approximate, e^(-64 x 0.75 / 32)), so either the base length scales with the register count (the verifier's 10 ms gate holds to about 330,000 ops) or the extra registers belong to the shadow alone and part (b) reads the base's 8 | modelled; the census decides |
Verdict: KEEP as layer 6, conditional on layer 5 (alone it moves nothing). Its value is one sentence in the chip's project plan: no interposer saves the second die.
### 2.3 Layer 7: warp-uniform data-dependent block selection (KEEP, rank 3, small)
A program whose control flow depends on the data it reads is the brief's first candidate. Per-lane branches are dead on arrival: a divergent branch costs a GPU warp both paths and a chip with per-lane sequencers nothing (the history's "placed nowhere" table; RandomX's one predictable branch targets speculative CPUs, which Igneum does not have). The form that survives is warp-uniform: at the end of each iteration a value reduced across the 32 lanes by shuffles (xor-fold of `r0`, say, which costs 5 shuffles) selects which of B drawn sub-blocks runs next, the same block for every lane of the warp, so the GPU takes one uniform indirect branch per iteration (as `sel` already takes one per iteration for the immediates) and the FPGA overlay or the fixed pipeline must hold all B blocks and pay B times the shadow's area for one block's throughput. For the GPU-like chip, a sequencer that already runs the hour's program, it is one more jump. What it buys: the per-epoch bitstream (the FPGA lane, history addition 5) holds B blocks instead of one, so a mid-size part's compile (42 to 160 minutes, PRflow, claimed in spec 1.13.1) carries B times the logic; at B = 4 a part that fitted one block does not fit, and at B = 8 the overlay must time-multiplex. Nothing in the energy identity moves. The acceptance must run every reachable path (B blocks per iteration, each judged by (a') in its own order) and the uniformity census must read the selection's bias (a selection that favours one block is a block that runs more).
Known-failed test, as a harness run: a research-only class `mx8+sh256x27+sel<B>` (the generator draws B blocks, the interpreter selects per iteration from the warp-folded `r0`), then `igneum-pow accept` over 256 seeds with every path judged, and `attack-f8 warps` at 2^24 on 64 seeds. Fail: the block-selection histogram over the 2^24 nonces outside 6 sigma of uniform (a plant: select from lane 0's `r0` bit 0 alone, which the fold is meant to prevent), or any path's (a') verdict differing from the whole-program verdict. Pass: within the band, 60 of 64 seeds under 1.2x on the hot-set test, as class v4 reads.
Chip row:
| Column | Value | Label |
|---|---|---|
| A fixed-function chip or an FPGA overlay | B x the shadow's logic for one block's throughput, or time-multiplexing at 1/B the rate; a bitstream compiled per epoch carries B blocks | approximate (LUT area scales with the straight-line block; no FPGA row exists in the repo) |
| A GPU-like chip | nothing: one jump per iteration on a sequencer | modelled |
| RTX 5090 / 5070 Ti / M5 Max | about 0: one uniform branch per iteration, 5 shuffles per iteration for the fold (5 x 8 = 40 shuffles per hash at 29 to 56 pJ: 0.002 microjoules, 0.3 percent of `F`); the compile-ahead carries B x 256 instructions of text (B = 4: the 1,024-line footprint that cost the M5 Max 17 percent on 6 October) | modelled on measured per-op costs; the Apple footprint is the cap on B |
| The verifier | B x the acceptance's dynamic test per candidate (every path); the hash's cost unchanged | modelled |
Verdict: KEEP, rank 3, with B capped by the Apple footprint (B = 2 or 4 at the 256-instruction block, or B = 4 at 64-instruction blocks, which the 6 October measurement says run 2.5 to 3.5 percent faster anyway). It is the only candidate that moves the FPGA lane, which the history ranks as the first adversary of a per-hour program (Lyra2REv2, X16R) and which no measured row in the repo has priced (the HBM FPGA row is 0.30x to 0.39x of a 5090 per watt, chip-model-v3 5.3, the soft-overlay case unmeasured).
### 2.4 The reserve ordered by hardware orthogonality (KEEP as an ordering rule inside layer 3)
Layer 3 unlocks reserve families by height and rotates after exhaustion. The order is Open in spec 1.13.2 except R1. The history's addition 6 said: families that force a full 32-bit datapath per lane first, the int8 tile last. Today's measured rows (research file 15.1a; the research lane's layer 3 table) sharpen it: a shuffle crossbar is the one block where the GPU's cost per op is highest (29 to 56 pJ) and a chip's is near it (about 20 pJ, approximate), so `shfla` (lane plus delta, a second crossbar form) is the family a 12-op chip lacks most and gains least on; byte permute and popcount next (small adders a chip adds for 0.1 pJ, but a datapath without them loses 4 points of the mix); the int8 tile last (k 0.03 to 0.3: a chip's MAC array is cheaper than the GPU's tensor core, so the tile is kept for datapath diversity and never for joules). The chip row is layer 3's: about USD 4 of N5 pre-wires all eight. Cost to the cards at 4 points: under 1 percent of ALU time on every vendor (measured step costs: shfla 1.91x on Apple, 1.53x NVIDIA, 0.75 to 0.84 AMD). Known-failed test: the layer 3 gate's own (a kernel built without the live family refused at packcheck; the fast-time harness crossing one family epoch with a stale miner, 0 accepted blocks after it). Verdict: not a new layer; an ordering rule, stated so the synthesis fixes it at genesis.
### 2.5 Data-dependent program graphs, per-lane (REJECT)
The brief's form: the program's control flow drawn from the data it reads, per lane. A GPU warp executes a divergent branch as both paths with lanes masked, so a branch taken by half the lanes doubles the ALU work of that span; a chip with a sequencer per lane pays the taken path only. Number: a shadow of 55,296 instructions per hash with one two-way branch per 64 instructions at 50 percent divergence costs the card up to 2x the shadow's premium (165 W instead of 83 at the knee on the 5090, modelled on the measured 6.4 pJ per op) for a chip cost of 1x; `k` on the branched work falls to 0.15 to 0.4. Costs GPUs more than chips. The warp-uniform form (2.3) is what survives.
### 2.6 Latency-bound reads tied to the shard proof (REJECT)
The hash's reads sampled from the state the miner is proving: class v5 already keys every item to a leaf of the execution state at the epoch's cut and refreshes per epoch (spec 1.8.6; proof of following). Tying the reads to the segment being proved means a refresh per block (every second) from the touched leaves. The research file's row 7 priced the refresh as a cost: a 1 GiB rebuild is 157 G ops, 13.4 ms on a 5090 and 32 ms on a 4090 (measured), 0.16 to 0.5 J on a chip core (1 to 3 pJ per op, approximate); per block that is 0.3 to 1 W against 78 W of chip hashing (0.4 to 1.3 percent) and 1.3 to 3.2 percent of a GPU's hash time (the rebuild stalls the hash on the card; the chip's rebuild runs on its core beside the memory). A delta refresh (only the touched leaves, a few KB) costs both sides nothing. Either way the GPU pays more or equal. What the tie would buy is liveness (a chip must follow the chain per block, not per epoch), which class v5's per-epoch refresh already gives at the WAN line of 2a.2. Number: GPU 1.3 to 3.2 percent of rate against a chip's 0.4 to 1.3 percent of energy. Rejected.
### 2.7 Randomised memory topology per era (REJECT)
The dataset's address map and stride family drawn per era: class v3 draws the stride multiplier `M`, the rotation `R` and the interleave `pos` per era already (spec 1.13.1, Counter ASIC 2.0 layers 4 and 8, decided IN at a six-era hash-rate spread of 1.3 percent on the 5090, 3.2 on the 9070 XT, 0.8 on the M5 Max, measured). The plan said then what still holds: a chip whose address decoder can permute its address lines pays nothing. Drawing a richer family (a per-era permutation polynomial over bank and row bits, a drawn item size, a drawn line interleave across devices) costs the chip's decoder a few hundred gates and costs the honest card whatever the mapping does to its own DRAM's bank parallelism: a mapping that concentrates consecutive dependent reads into one bank group hurts the side with fewer lanes in flight, which is the chip (1,172 against 7,262), but the chip adds lanes at 64 B each (lane state is free, chip-model-v3 5.5), so the asymmetry closes at no cost. The one topology lever that would have moved the chip, the hot region above the window model, was read by adv-cache-2 as the diffuse era-stride excess (a product's low bits placed at address bit `R`; 8 of 27 drawn-era programs over 1.2x), which is a FAULT the next class's value-level test removes, not a lever to keep. Number: 0 to the chip, 0 to 3.2 percent to the cards. Rejected; the existing draws stand.
### 2.8 VRAM-size ratchet (REJECT as a lever; layer 2's floor stands)
A floor that rises with chain state by rule is layer 2. The ratchet form (the floor tracking the modal miner's VRAM minus the prover footprint, or rising on a calendar faster than the schedule) was read against the card-lifetime table (`docs/analysis/card-lifetime-2026-10-05.md`, option (b) steps): the 4 GB tier ends at the 4 GiB step, 8 GB at 8 GiB, 12 and 16 GB at the 16 GiB step; a chip holds 24 GB (one HBM3 stack, about USD 200, modelled) or 32 GB (the 5090's own 16 devices, USD 320), so every step retires a card tier before it touches the chip, and at 32 GiB and beyond the chip adds devices and its activate-bound rate RISES with the bank count (chip-model-v3 5.7, row "dataset size": "not a lever against this chip"). Number: at the 16 GiB step the 8, 12 and 16 GB tiers are out (3 of 6 card tiers) and the chip's energy per hash moves 0. Rejected as a lever; layer 2's rule (the schedule as the floor, the state above it, a ceiling at the next cache doubling) is kept exactly as the synthesis writes it, with its honest line that it retires cards before chips.
### 2.9 Proof-carrying hashes sampled by the pool (REJECT)
A fraction of hashes carries a verifiable execution witness. Three witness forms were read. (i) The hash's own 128 item values: the stored-dataset chip has every item; the recompute chip derives them; cost 0 to both, 512 B per share on the wire. (ii) A Merkle witness of the state leaves under the window's root: the chip's node has it (one node serves a farm, class-v5 2a.2); cost 0 to both. (iii) A witness that the item was DERIVED (a transcript of the 8 dependent cache reads and the mixer's 72 applications): a stored-dataset chip cannot produce it without the cache and the mixer core, so this form forces the `f = 0` chip's silicon (the 256 MiB SRAM mirror, USD 46 of die and an N5 project) onto the `f = 1` chip for the sampled fraction `g`; but the honest GPU must produce the same transcript, and deriving an item on the card is 8 dependent 64-byte cache reads at the mixer's 9,360 ops (the inline kernel measured 4.8x slower than the honest kernel on the M5 Max, spec 1.8.5), so at `g = 1/128` (one item per hash) the card pays about 4 percent of its rate and at `g = 1/16` about 30 percent; the chip derives on an SRAM-resident cache at 6.3 nJ per item (modelled) for 7 percent of its energy at `g = 1/16`. Number: GPU 4 to 30 percent of rate against the chip's 1 to 7 percent of energy. Rejected on form (iii); forms (i) and (ii) force nothing.
### 2.10 Prover-gated eligibility (REJECT)
The block's eligibility tied to the miner's proving (a key must have proved its share of segments in the last window to claim a block), so a chip farm must carry provers. The bound is the gas bound the research file's section 7 found: the chain needs 2 shards per block at the v1 budget, about 1.6 kW of 5090 proving network-wide at 1 block per second (measured prover rows), independent of the hash rate. Against the hash: at 1 GH/s the hash draws 2.4 kW (2.4 microjoules per hash, measured), so proving is 67 percent of it; at 100 GH/s 0.7 percent; at 10 TH/s 0.007 percent. A chip farm at share `s` must prove share `s` of 1.6 kW: six 5090s per farm at any scale, which is the node it already runs. Redundant proving (each segment proved by `m` provers) raises the forcing `m` times and is the useful-work gaming the history records (Aleo, Boundless). Number: at mainnet scale the forcing is under 0.01 percent of the chip's energy; the honest card already proves. Rejected; the 80/20 split stands.
### 2.11 Time-locked parameter commitments (REJECT beyond the era VDF)
An era's parameters committed under a VDF so a chip cannot be built ahead: the era VDF of 7 October (`docs/analysis/era-vdf-2026-10-07.md`) already makes the era draw's input unknowable for 517 s on the fastest prover measured (chiavdf's GMP path, 208,800 squarings per second, against the production T of 108 million) and the era lead is 2 hours (`pow_era_lead`), so a chip taped out against era `n` knows era `n + 1`'s draw 2 hours before it runs, against a 5-month (Bitmain) to 13-month (a startup) design cycle (the history's lesson 5, Vorick). Lengthening the delay or the commitment changes nothing a chip can use: the GPU-like chip holds the whole drawn band as firmware (the synthesis's section 7), and the fixed-function chip is dead at the first draw outside its wired value whether it learns the draw 2 hours or 2 days ahead. The one party for whom 2 hours matters is the FPGA fleet (a bitstream compiles in 42 to 160 minutes on a mid-size part, claimed), and layer 7 (2.3) and the epoch length (spec 1.13.1, the 600 s floor) are the levers for it, not the lock. Cost of a longer lock: one honest node core for the VDF's hour per era (today) rising linearly with T; the 2019-class verify gate already missed by 2.2x (26 ms against 10, measured, proxy). Number: 0 to the chip at any delay above 2 hours; the honest node's core-hours rise with T. Rejected.
### 2.12 A fraction of reads derived from the cache in the hash (REJECT)
The brief's spirit of "tie the hash to what the chip must hold": a fraction `g` of the 128 reads per hash derived on the fly from the 256 MiB cache (8 dependent cache reads and 72 mixer applications) instead of read from the dataset, so the stored-dataset chip must carry the recompute chip's cache and core for that fraction. This is 2.9 form (iii) without the witness and the same arithmetic: the card's derived read is 8 dependent DRAM reads (the cache does not fit L2 at 256 MiB, and the cache doubling keeps it so), so at `g = 1/16` the card's dependent-read count per hash rises from 128 to 184 and its rate falls about 30 percent (modelled on the latency-bound rule; the inline kernel's 4.8x at `g = 1` is the measured anchor); the chip with the cache on die derives at 6.3 nJ per item (4.0 nJ of SRAM reads, 2.3 of mixer; modelled) for 0.466 to 0.50 microjoules per hash (+7 percent) and buys the USD 46 mirror and the N5 project it already needs for the shadow core. Number: GPU -30 percent of rate at `g = 1/16`, chip +7 percent of energy and +USD 46 of die. Rejected.
### 2.13 Row-straddling and double-activation reads (REJECT)
The chip and the card share the DRAM's physics (the research file's section 5: the same tRC, the same activate window, the same 32-byte atom). A read that opens two rows (an item straddling a row boundary, or two independent 32-byte sectors per read) costs the chip's memory +0.9 nJ per read (a second 909 pJ activation; modelled) and the card +1 sector of traffic, which at 128 reads per hash is the w64 regime (64 B per read): the 5090 fell to 71.9 MH/s, bandwidth-bound, 47 percent of its rate (measured, read-width). Number: chip +45 percent of `E_mem` (0.466 to 0.58), card -47 percent of rate and about +9 percent of energy per hash on the memory side alone. The edge moves from 3.6x to about 3.1x at the knee (modelled) at the price of half the card's rate. Rejected.
### 2.14 Per-era lane-state and scratch draws (REJECT)
Per-lane live state across the hash (a scratch with read-modify-write) was measured out in Counter ASIC 2.0 (layer 3: the recompute chip's gain at every share 2.4x, the cards -12 to -48 percent) and bounded in `docs/analysis/scratch-soundness.md` (the live state sits in a chip's SRAM at under 5 percent of its mirror). A drawn scratch size per era draws from a dead family. Number: the cards -12 to -48 percent of rate (measured), the chip +picojoules per access. Rejected. (The register-file width of 2.2 is the live form of this idea: state that costs the GPU nothing because its register file is already there, and costs the chip a link, not an SRAM.)
### 2.15 A refresh per block (REJECT; the research file's row 7)
Dead by arithmetic: a 1 GiB rebuild is 0.16 to 0.5 J on a chip core and 13 to 32 ms of a card's hash time; per block that is 0.4 to 1.3 percent of the chip's energy and 1.3 to 3.2 percent of the card's rate. The refresh cadence is a liveness tool (class v5's proof of following), not an energy lever.
## 3. The measured rows (igneum-build-1, 8 October 2026, 11:0x to 11:2x UK)
The crate: `igneum-pow` of branch `counter-asic-4` at 5984ffab, built on build-1 through `tools/build-remote.sh --no-fetch --box 1` from the detached worktree `igneum-wt-v6-inv-ca4` (RESULT rc=0, 11 s, sccache); the binary `/srv/builds/igneum-wt-v6-inv-ca4/igneum-pow/target/release/igneum-pow`. Every run under `/srv/builds/_bin/lease` (pool 48 at class measure for the censuses, `cores 40,88` for the benches), owner `class-v6-invention`; the box at load 15 to 24 on 96 threads from other lanes throughout; logs under `/srv/builds/v6-invention/` (`census-none.tsv`, `census-era.tsv`, `logs/<era>-<form>-<seed>.log`, `census-run-*.log`), copied into `docs/analysis/class-v6/logs/` with the 09:00 report.
### 3.1 The acceptance census (`v6inv-census.sh`, 256 seeds `igneum-v6inv/0..255`, every candidate's verdict through `igneum-pow accept --class <form>`, the class's own 32-attempt cap)
No era (`ERA=none`; 2,816 rows in 80 s on 48 cores):
| Form (sub-blocks x length x passes) | Shadow instructions per iteration | Seeds accepted of 256 | Seeds exhausting 32 attempts | Candidates | Rejection per candidate | Mean accepted attempt | First failing part, the top four |
|---|---|---|---|---|---|---|---|
| `mx8+sh256x27` (class v4's shape, the whole block after instruction 63) | 6,912 | 256 | 0 | 268 | 0.045 | 0.05 | (a) 6, (b) 5, (c) 1 (the v2/v3 rule only: the crate's `accept` does not take the class v4 parts on this spelling, so the row is the control's shape, not sub-version 3's 0.681) |
| `mx8+shl256x27` (16 x 16 x 27, the 7 October form) | 6,912 | 76 | 180 | 7,043 | 0.989 | 15.9 | per-load bias 6,081; (c) 545; (b) 212; (a) 129 |
| `mx8+shl256x16` (16 x 16 x 16) | 4,096 | 69 | 187 | 7,117 | 0.990 | 15.4 | bias 6,294; (c) 406; (b) 215; (a) 133 |
| `mx8+shl576x12` (16 x 36 x 12) | 6,912 | 215 | 41 | 3,815 | 0.944 | 10.6 | bias 3,367; (b) 98; (a) 72; (c) 63 |
| `mx8+shl512x8` (16 x 32 x 8) | 4,096 | 209 | 47 | 4,120 | 0.949 | 11.5 | bias 3,686; (b) 106; (a) 82; (c) 37 |
| `mx8+shl1152x6` (16 x 72 x 6) | 6,912 | 226 | 30 | 3,428 | 0.934 | 9.9 | bias 3,036; (b) 97; (a) 62; (c) 7 |
| `mx8+shl1024x4` (16 x 64 x 4) | 4,096 | 231 | 25 | 3,487 | 0.934 | 10.6 | bias 3,097; (b) 81; (a) 73; (c) 5 |
| `mx8+shl2304x3` (16 x 144 x 3, class v4's count) | 6,912 | 224 | 32 | 3,457 | 0.935 | 9.9 | bias 3,076; (b) 89; (a) 67; (c) 1 |
| `mx8+shl2048x2` (16 x 128 x 2) | 4,096 | 228 | 28 | 3,423 | 0.933 | 10.1 | bias 3,041; (b) 83; (a) 65; (c) 6 |
| **`mx8+shl4096x1` (16 x 256 x 1, the sound form inside the cap)** | 4,096 | **234** | **22** | 3,217 | **0.927** | 9.7 | bias 2,823; (b) 94; (a) 64; (c) 2 |
| `mx8+shl4096x2` (16 x 256 x 2) | 8,192 | 233 | 23 | 3,273 | 0.929 | 9.9 | bias 2,875; (b) 95; (a) 67; (c) 3 |
What the ladder says: the iterated 16-instruction map is the fault (0.989 to 0.990 whatever its pass count), and from a 36-instruction sub-block up the rejection is flat at 0.927 to 0.949 whatever the pass count or the work per hash (4,096 or 6,912 or 8,192 instructions per iteration). The remaining 0.93 is not the placement: the `(c)` lane-constant and distinct rejections fall to 1 to 7 per form (they were 406 to 545 on the 16-instruction forms), and the bias rejections are the product's low-bit law at the load's source, which under class v4 the sub-version 3 dataflow rule (a') removes at the draw and which the per-load class, not being the class v4 shape, never applies. The value-level reading of the 256 x 1 form's 2,329 bias rejections: index bit 0 in 1,563 (one-counts clustered at 3,584 to 4,608 of 16,384, the 1/4 law, 680 of them; and at 11,776 to 12,288, the 3/4 complement, 212), bits 26 and 27 in 502 (one-counts 7,168 to 7,680: a mild low bias of the top address bits just outside the 6-sigma band of 384, the high bits of small products through `mulhi`, unattributed), the other 26 bits 264 in all; by site, site 0 takes 788 of 2,329 (its source is written last by the previous iteration's sub-block 15 and by the base instructions before instruction 1, where the draw's redraw covers the sub-block's last writer of the next load's source and not a product rotated into place by a later `rotl`). The named fix is one rule, the research file's own ask of 20.2: the sub-version 3 freshness fixpoint and the shared-operand rule run over the real execution order (base instruction, the load, its sub-block, the next base instructions), and a load whose source is not fresh at that point refused at the draw, which class v4 pays at 0.568 of candidates and which should bring the per-load forms from 0.93 toward 0.68.
The acceptance's own cost: 35 ms per per-load candidate on one box core (1,111 ms for 32 candidates; the control's v2/v3 rule 2.8 ms), so an epoch's draw at 13.7 candidates is about 0.5 s of one core; the (c'') ratio at 2^20 would add class v4's 2.8 s per chosen candidate.
### 3.2 The verifier (core 40 of the EPYC 9454P at nice 19, core 88 its SMT sibling held busy by a 100,000-warp class v4 bench for the whole run, killed by its pid at the end; the ladder's method; `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <form> --warps 50`)
| Form | Cold, sibling loaded, warp base 0 / 4,096 / 1,000,000 (ms) | Avg of 50, loaded (ms) | Cold, core alone, earlier in the same minutes (ms) | Label |
|---|---|---|---|---|
| `mx8+sh256x27` (class v4's shape) | 8.63 / 8.33 / 8.33 | 8.29 | 5.90 / 6.34 / 5.24 | measured; the ladder's run 2 read 8.77 / 5.14 at load 25 on 6 October |
| `mx8+shl4096x1` | 8.55 / 8.28 / 8.28 | 8.24 | 5.44 / 5.21 / 5.17 | measured |
| `mx8+shl2304x3` | 8.82 / 8.54 / 8.44 | 8.43 | 5.31 / 5.13 / 5.09 | measured |
| `mx8+shl4096x2` | not run loaded | | 5.97 / 5.33 / 5.29 | measured, alone only |
The per-load forms verify in class v4's time within the run's noise (the same instruction count, the same items derived, 4,096 per warp on every row). A first pass of the loaded column was discarded: its sibling run (300 warps) ended before the measured warps began, the ladder script's own known-failed case, and read the quiet-core figures; the second pass held the sibling for the whole run.
### 3.3 The instrument fault under drawn eras, and the correction it forces
The same 2,816 rows with `--era igneum-era-test/<seed mod 16>` (`ERA=era`): every per-load form 0 of 256 seeds accepted, 8,192 candidates per form, rejection 1.000; the class v4 shape 256 of 256. Of the 256 x 1 form's 7,862 bias rejections, 5,536 name index bit 26 or 27 "set in 0 of 16,384" or "set in 16,384 of 16,384" at a load site, which is the era's working-set window (spec 1.13.1: `k_off = below(3)` per site puts the site on the whole dataset, a half or a quarter by fixing the top `k` bits of the index to the drawn offset `o`); the prototype's `BiasedIndexBit` test loops bits 0 to 27 and does not exclude bits at or above `28 - k_off_s`, so under any era it refuses every program with a half- or quarter-window site, which is nearly every program. The research file's 20.2a-close ("the 16 x 27 per-load class accepts 22 of 1,621 candidates over 64 seeds, 1.4 percent; 42 of 64 seeds exhaust the chain's 32 attempts") was read across drawn eras and therefore measured the instrument on most of its rows; the no-era census above is the construction's own figure (the 16 x 27 form 0.989 per candidate, 76 of 256 seeds accepted, which still fails the test of 2.1 and keeps that form dead). Owed to the research file (the Counter ASIC coordinator): the instrument fix (skip the window bits per site) and the 20.2a-close figures re-read with it; neither is made this week by this lane, which changes no code in the crate.
### 3.4 The packs
`igneum-pow export --seed igneum-genesis --day 2026-10-03 --class <form> --out <dir>` on build-1: `mx8_shl4096x1` (attempt 7, id 75ca9547da21b200, OVERALL PASS, kernel.cu 245,267 bytes, 4,308 lines) and `mx8_shl2304x3` (attempt 8, id bbfdfc1dcdda0b46, OVERALL PASS, kernel.cu 142,010 bytes, 2,516 lines), under `/srv/builds/v6-invention/packs/`, tarred as `v6inv-perload-packs.tgz` (sha256 ca1986b7fc5fab20a643fc37151a55e01f91edbfacc6f1a22a7384ae87cc11bc). Handed to the hash lane at 11:1x UK, taken into its v6 PC 1 job (the 5090 alone, the packs beside today's class v3 x8 and x16, the two re-weighted shadow packs, and the controls `mx8-genesis` and `mx8_sh256x27`; self-test, the 2^24 fingerprint, MH/s and W unlocked and at the 1,300 MHz lock), queued after the AMD grid, publish about 12:00 UK, close about 12:45 UK. Nothing of the packs' output goes to a served page or the spec.
## 4. The ranked list, every column
| Rank | Layer | Door (section 1) | Chip: fixed-function `k` / capex | Chip: GPU-like per-joule edge | Cost: RTX 5090 | Cost: RTX 5070 Ti | Cost: Apple M5 Max | Verifier | Harness state |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Layer 5: the shadow per load, one pass of a long sub-block | capex and project (door 2) | `k` unchanged; project about USD 60 M against 30 M, cap about 200 M against 100 M (modelled); capex 3.0 to 4.3 USD per MH/s | unchanged (2.1x at `k = 1`) | -13 to -14 W against class v4's block at the same work (16 x 27 form, measured); the sound forms' rows in today's PC 1 job | no row; the 4070's 30 W premium as the proxy (approximate) | -0.5 percent (16 x 27 form, measured); the 4,096-line footprint OWED | 8.28 to 8.82 ms loaded (measured) | acceptance 0.927 measured; the dataflow rule in execution order is the fix (gate plan) |
| 2 | Layer 6: the register-file width drawn per era | capex (door 2), with layer 5 | closes the interposer branch: 4.5 to 9 TB/s of die-to-die traffic (modelled, approximate link figures) | unchanged | about 0 rate (occupancy arithmetic, approximate; the pack measures it) | the same | the same | 4x the register arrays, cost per op unchanged (modelled) | no pack form; a generator constant (gate plan) |
| 3 | Layer 7: warp-uniform block selection | life of a bitstream (door 3, the FPGA lane) | B x the shadow's area for a fixed pipeline or an overlay (approximate) | unchanged | about 0 (one uniform branch and 40 shuffles per hash: 0.3 percent of `F`) | the same | the compile footprint caps B (the 1,024-line block cost 17 percent, measured) | B x the dynamic test per candidate | no pack form (gate plan) |
| 4 | The reserve's order | life of fixed silicon (door 3) | USD 4 pre-wired | unchanged | under 1 percent at 4 points (measured steps) | the same | the same | 0 | layer 3's gate |
What the list does not contain, and why it is honest to say so: no candidate moves the per-joule identity. The chip that stores the dataset keeps 3.6x at zero premium and 2.1x at `k = 1` under every layer here, as under the four. What the list adds is the project cost (rank 1 doubles it on the mission lane's model, rank 2 removes its cheaper branch), the FPGA lane's cost (rank 3), and the order of the reserve (rank 4). The founder's line "render an ASIC useless as soon as it dropped" is true of the fixed-function chip under layers 1 and 3 and stays false of the GPU-like chip under everything; the honest public sentence is the research file's: the price per joule of the honest card's operating point and the shadow's premium are what hold the general chip, and the layers decide which chip can be built and what it costs to start.
## 5. Consequences per tier (the standing rule of 5 October 2026)
| Tier | What this file means today | What is being done |
|---|---|---|
| A home miner, one 8, 12 or 16 GB card, any vendor, any OS | nothing changes: no layer here ships this week, no class moves, the devnet pays nothing; if rank 1 lands in v6 the card runs the same shadow work in a different place at the same or lower watts (the 16 x 27 form's measured 13 to 14 W under the whole block on a 5090; a 4070-class card's premium at its tune point is 30 W today, measured) | the sound forms' 5090 rows today; a small-card row (the 4070) by job when the hash lane's queue allows |
| One 24 or 32 GB card (5090 class) | the measured rows of section 3; the per-load placement is cheaper for this card than class v4's block at the same work | the PC 1 job, about 12:45 UK |
| RTX 5070 Ti | no card in the project; every row is the 4070's or the 5090's scaled, approximate; the hash lane's default (the stock pair plus the 5090's lock slope) stands | stated as approximate wherever it appears |
| Apple (M-series) | the one open number: the inline footprint of a 4,096-line per-load block (the 1,024-line block cost 17 percent on 6 October); if it costs rate, rank 1 takes a block-shape cap for the Apple tier (144 x 3 at 2,516 lines, or a 64-instruction sub-block form) and the synthesis says so | a Metal packbench row under the measure lock by the lane that runs the Mac (not this one) |
| A rig | watts per card as the 5090 row; a rig's bill under rank 1 is at or under class v4's | the same rows |
| A pool user | nothing: no share, payout or template changes in any candidate kept; the rejected 2.9 (pool-sampled witnesses) is the only one that would have touched the pool protocol | nothing |
| A node operator (the verifier) | the per-load forms verify in class v4's time (8.3 to 8.8 ms loaded, measured); layer 7 at B blocks multiplies the acceptance's dynamic test per candidate, not the hash | the 2019-class core measurement (O-1.14) decides rung and block caps as before |
| A chip maker | under rank 1 the controller and the shadow core share one advanced die or an interposer (project about USD 60 M, modelled); under rank 2 the interposer no longer suffices; under rank 3 a per-epoch bitstream carries B blocks; the per-joule edge is unchanged | the gate plan of section 6 |
| The public claim | nothing moves; the chip texts rest on the research file's close (2.1x at `k = 1` for 82 to 90 W on a 5090 at the knee, measured four times) | the Counter lane's texts |
## 6. The gate plan (hours, never weeks; nothing this week)
| Gate | What runs | Pass line | Known-failed case |
|---|---|---|---|
| G5-draw (layer 5) | the generator's dataflow fixpoint and shared-operand rule over the real per-load order; the 4,096 cap raised so 16 x 432 x 1 draws; `v6inv-census.sh` at 256 seeds, no era and drawn eras (with the instrument's window bits excluded) | rejection per candidate at or under 0.75 on every sound form, 0 exhaustions, drawn-era rate equal to the no-era rate within the binomial band; the F8 census at 2^24 on 64 seeds: 60 of 64 under 1.2x, the four tail seeds' sites read against their own windows | the 7 October per-load record (candidate 0 of 854050a4293f0615) refused; the 16 x 27 form at 0.989 refused by the line |
| G5-card (layer 5) | one pack per sound form on the 5090 (today), the 4070 and the 9070 XT by job, Metal packbench on the M5 Max | rate within 1 percent of class v4's shape at both states on NVIDIA and AMD; watts at or under class v4's; the Apple footprint within 5 percent or the block-shape cap set | a pack whose fingerprint differs from the Rust verifier's on any vendor |
| G6 (layer 6) | the register count as a class field; the census at 8, 16, 32; the card packs | 2.2's pass line | 2.2's fail line |
| G7 (layer 7) | the selection class; every path judged; the selection histogram at 2^24 | 2.3's pass line | the lane-0 plant |
| G-order (the reserve) | layer 3's gate with the order fixed | layer 3's line | layer 3's case |
Hours: the dataflow rule in execution order 2 to 3 (the research file's own estimate) plus the re-export and census 1; the register class 3 to 4; the selection class 4 to 6; the card jobs are queue time. Nothing is coded this week; the first code is the founder's call after the synthesis.
## 7. Unverified and owed
- The sound forms' card rows (the 5090 today through the hash lane; the 4070 and the 9070 XT later; the M5 Max footprint by the lane that runs the Mac).
- The 16 x 432 x 1 form itself: bracketed by the 256 x 1 and 144 x 3 rows (0.927 and 0.935) and the flat ladder between 36 and 256; not drawn (the cap).
- The instrument fix (the window bits) and the re-read of 20.2a-close: owed to the research file's owner, not made here.
- The bits 26 and 27 mild bias (502 rejections at one-counts 7,168 to 7,680 of 16,384): unattributed; a trace of the site's source writers is the next read.
- The register-width and block-selection classes: modelled only; no pack exists.
- Every chip figure is the model's (chip-model-v3 section 5 and the research file's sections 2, 15.1a, 16 and 20.4); the die-to-die link figures of 2.2 are approximate, from memory, uncited; no chip has been measured.
- The web search budget of this session was exhausted before this lane's reading; every external figure here is one already cited in the repo's files, with its URL and date there (the research file's section 14, the history's section 6, chip-model-v3 5.1).
## 8. Sources
Internal: `docs/spec/01-lottery-hash.md` (1.4.3 to 1.4.7, 1.8.5, 1.8.6 on branch class-v5, 1.13); `docs/analysis/chip-model-v3.md` (sections 1 to 3, 5.1 to 5.11, 6); `docs/analysis/counter-asic-4-research.md` on branch counter-asic-4 (sections 0, 2, 4, 7, 9, 15.1a, 15.1b, 16, 17, 20.2 to 20.4); `docs/plans/cryptanalysis/in-house-pass.md` (sections 12 to 14); `docs/plans/counter-asic-3-status.md` section 7c; `docs/analysis/asic-resistance-history.md` (sections 1.2, 2.4 to 2.6, 3, 4.3); `docs/analysis/era-vdf-2026-10-07.md`; `docs/design/class-v5-stored-state.md` on branch class-v5 (2a, 3); `docs/design/latency-ladder.md` on branch ladder (2 to 5); `docs/analysis/card-lifetime-2026-10-05.md`; `docs/analysis/latency-shadow-2026-10-06.md`; `docs/analysis/scratch-soundness.md`; `docs/design/class-v6-rotating-family.md` on branch counter-asic-4 (the synthesis's outline at 5984ffab); the harness runs of section 3 (logs on build-1 under `/srv/builds/v6-invention/`).
External, as cited in those files (read there on the dates they state): O'Connor et al., Fine-Grained DRAM, MICRO 2017; Li, Reddy, Jacob, MEMSYS 2018; Folded Banks, ISCA 2025; Horowitz, ISSCC 2014; Dally, Hot Chips 2023; the mlsysbook energy table citing Horowitz and Dally; the RandomX design document; the ProgPoW audits (Least Authority and Rao, 2019); Condrey, PoSME, arXiv 2604.15751 (April 2026); the Ethash, RandomX, Equihash and Cuckatoo chip rows of the history; the GDDR7 price rows (TrendForce, September 2026).

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -288,7 +288,7 @@ the height arrives.
(26640/28640, a follower, lowest risk), the seed (`/opt/igneum/v4/bin`, `infra/seed-nodes/stage-v4.sh`),
Mac node 1 (26610/26611), PC 1's node (its launcher's `igneumd.exe`; PC 2 mines through the Mac node
and needs nothing). Each node gets the same `"difficulty_v2_activation_daa": N` in its override file.
the project lead restarts the live processes; this entry does not.
The founder restarts the live processes; this entry does not.
3. Watch the observer's difficulty events across the height and the next epoch boundary; with a miner
joining inside an epoch the floor and the bursts of section 1 must not return.

View file

@ -64,7 +64,7 @@ The change, if the finality lane wants the certified binding: `EraVdfManager::cu
## 5. The verify gate on a 2019-class core
The gate was "the VDF verifies in under 10 ms on a 2019-class core". Measured: 21.9 and 22.6 ms with the group held on the box core (above), which the F6 row's calibration puts at about 1.19x on an i7-9700K (the O-1.14 run: 6.0 ms on the 2019 core against 5.06 ms on the box proxy), so about 26 ms on a 2019 core, labelled a proxy: no 2019 host was rented this lane (Vast rentals are a purchase; not made without the project lead's word). NOT MET, by 2.2x on the box and about 2.6x on the proxy.
The gate was "the VDF verifies in under 10 ms on a 2019-class core". Measured: 21.9 and 22.6 ms with the group held on the box core (above), which the F6 row's calibration puts at about 1.19x on an i7-9700K (the O-1.14 run: 6.0 ms on the 2019 core against 5.06 ms on the box proxy), so about 26 ms on a 2019 core, labelled a proxy: no 2019 host was rented this lane (Vast rentals are a purchase; not made without the founder's word). NOT MET, by 2.2x on the box and about 2.6x on the proxy.
Where the time goes and what closes it: a verification is two 256-bit exponentiations, about 770 group operations at 28 µs each; the operation is NUDUPL on the fixed-width integer, whose cost is the extended gcd (Lehmer rounds on 8-limb numbers) and the reduction. Three rounds of this lane moved it from 345 µs (a Lehmer convention defect that fell back to plain division every round) to 28 µs (the convention, i64 word division, 34 limbs, the x86-64 128-by-64 division); the next 2.2x is chiavdf's Pulmark reducer (reduce only when `a` exceeds 8 limbs, O-4.6) and a limb-level NUDUPL that keeps the partial gcd's intermediates in words, or GMP through `rug` behind a feature on the x86-64 Linux and Windows builds (chiavdf's 208 K/s is 6x this evaluator, which would put the verification near 4 ms as the prototype measured), with the fixed-width path the fallback for wasm and macOS. Owed, not blocking: the verification runs once per era (180 days) on a node that imports a record rather than evaluating; every mining node evaluates and never verifies.

View file

@ -1,6 +1,6 @@
# Per-identity hash rate "decay" on the RTX 5090: diagnosis and fix
3 October 2026, miner-community-lead. Source data: the uploaded logs of the project lead's PC (`node tools/logs.mjs
3 October 2026, miner-community-lead. Source data: the uploaded logs of the founder's PC (`node tools/logs.mjs
nvidia-DESKTOP-KMCV30N-1-20261003-222331 --all` and the other identities, the launcher log
`igneum-DESKTOP-KMCV30N-20261003-222331`), the three serve loops, the miner's worker mode, and five runs of the
Metal worker on the Mac against private test nodes (ports 27500 and up, `/tmp/igneum-decay-test`). Figures
@ -9,7 +9,7 @@ from the logs are exact; the two labelled approximate are from memory.
## 1. Finding in one paragraph
There is no per-job growth in any worker or in the miner's memory. Two separate things produce the picture
the project lead saw. First, the STATUS line's two rates are cumulative averages since the miner started
The founder saw. First, the STATUS line's two rates are cumulative averages since the miner started
(`hashes_total / elapsed` and `hashes_total / gpu_ms_total` in `mine_worker`), so a fast first interval
decays as 1/t by construction; nvidia-1 was alone on the card for its first seconds and every later interval
ran at a flat 17.8 MH/s wall, while the last-started identity, nvidia-8, shows the mirror image, a cumulative
@ -28,7 +28,7 @@ The miner prints `hash=A MH/s wall (B MH/s inside jobs)` with `A = hashes_total
`dH / dt` and `dH / dG` with `H = A x t` and `G = H / B`. Every table below is that calculation
(`rates.py` in the bench-log entry).
### 2.1 The segment the project lead quoted: 22:57 to 23:04 UTC, after the epoch-3 restart (DAA 10,801 on)
### 2.1 The segment the founder quoted: 22:57 to 23:04 UTC, after the epoch-3 restart (DAA 10,801 on)
nvidia-1, jobs of 2^24 nonces, STATUS every 30 s:
@ -95,7 +95,7 @@ card going idle, which is what a growing CPU-side gap in every worker does.
## 3. Code audit: what is allocated per job, and what is freed
### 3.1 `proto-cuda/host.cu`, `runServe` (HEAD, the binary the project lead ran, built by the launcher at 22:09:57)
### 3.1 `proto-cuda/host.cu`, `runServe` (HEAD, the binary the founder ran, built by the launcher at 22:09:57)
| Allocation | When | Size | Freed |
|---|---|---|---|
@ -113,7 +113,7 @@ version (the hot-swap agent's) adds `CudaPair` (at most two resident, the old on
job on the new pair, `releasePair` frees dataset, cache and both modules) and `PrepareTask` (deleted after
the load); still nothing per job. `cudaDeviceSynchronize` per dispatch under the default
`cudaDeviceScheduleAuto` spins the host thread when the process holds fewer contexts than the machine has
cores, which is always true here (one context per process): that is the 6.2% CPU per worker the project lead saw (one
cores, which is always true here (one context per process): that is the 6.2% CPU per worker the founder saw (one
of 16 threads). The hot-swap working tree sets `cudaSetDeviceFlags(cudaDeviceScheduleBlockingSync)` before
the context is created (host.cu, main), which is the right call and the right place; the thread then sleeps
on the dispatch. It does not change the hash rate directly, but eight spinning threads plus eight OpenCL
@ -250,7 +250,7 @@ because it is measured from the template fetch, which precedes the walk).
## 5. The time-slicing hypothesis: one worker versus eight, at two fixed difficulties
the project lead's Task Manager reading (GPU memory flat at 14.7 GB, the card 99% busy, each CUDA worker at 6.2% CPU) and the
The founder's Task Manager reading (GPU memory flat at 14.7 GB, the card 99% busy, each CUDA worker at 6.2% CPU) and the
hypothesis that eight contexts time-slicing one card with "longer jobs as blocks get rarer" explain the decay. A
job is a fixed 2^24 nonces, so its length does not depend on the target, but the four runs below test the
hypothesis as stated: one Metal worker and eight, each at a fixed low difficulty (2^25, 0.25 founds per job)
@ -483,20 +483,20 @@ warnings; also saved as `docs/analysis/hashrate-decay-2026-10-03.patch`).
acceptance figure is table 2.2 flattening: the gap per job no better than 0.10 s at DAA 10,800 (3,600
into an epoch) with eight identities.
## 7. Two things for the project lead to check on the PC
## 7. Two things for the founder to check on the PC
1. Dedicated GPU memory over time. Task Manager, Performance, GPU 0, "Dedicated GPU memory usage", or
`nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu,temperature.gpu,clocks.sm,clocks.mem,power.draw,clocks_throttle_reasons.active --format=csv -l 10 > gpu.csv`.
Expected: flat from the moment the eighth worker prints `ready` (each CUDA worker holds 1 GiB dataset +
256 MiB cache + 32 MiB output + the context, about 1.6 GiB; eight of them 13 to 15 GiB of the 32 GiB,
which matches the flat 14.7 GB the project lead read). A line that climbs while the hash rate falls would mean a leak
which matches the flat 14.7 GB the founder read). A line that climbs while the hash rate falls would mean a leak
in the worker; none exists in the code and the Mac RSS traces are flat.
2. GDDR7 memory junction temperature and the throttle reasons. HWiNFO64, Sensors, under the GPU: "GPU
Memory Junction Temperature", "GPU Thermal Limit", "GPU Power Limit", "GPU Reliability Voltage Limit"
(each "Yes" or "No"), "GPU Effective Clock" and "GPU Memory Clock"; or the `clocks_throttle_reasons.active`
column above (`0x0000000000000004` is SW power cap, `0x0000000000000020` SW thermal slowdown,
`0x0000000000000040` HW thermal slowdown, `0x0000000000000080` HW power brake). Approximate thresholds,
from memory: the core starts to pull clocks around 83 C (the project lead's 55 to 59 C is far below it); GDDR6X on the
from memory: the core starts to pull clocks around 83 C (the founder's 55 to 59 C is far below it); GDDR6X on the
previous generations throttles from about 95 C junction and the hard limit is 105 C; GDDR7 figures are not
published, so treat anything above 90 C junction as the zone to watch and a "Yes" on any limit row as the
signal. The decisive sign of throttling is "GPU Effective Clock" falling while utilization stays at 99%;

View file

@ -1,8 +1,8 @@
# Horizon, October 2026: the ranked research across every system
Written 6 to 7 October 2026 by the Horizon coordinator (branch `horizon`, worktree `igneum-wt-horizon`) on the project lead's ask of 6 October 2026, 20:4x UK: "deep backward and forward predictive research and modelling for our algo and all of our systems: is there room for improvement, room for more coin utility, algo improvements, security improvements, anything we can do to make a 51% attack impossible, basically creating a level of polish that has not been seen before." Later the same evening: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", "be revolutionary", and "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
Written 6 to 7 October 2026 by the Horizon coordinator (branch `horizon`, worktree `igneum-wt-horizon`) on the founder's ask of 6 October 2026, 20:4x UK: "deep backward and forward predictive research and modelling for our algo and all of our systems: is there room for improvement, room for more coin utility, algo improvements, security improvements, anything we can do to make a 51% attack impossible, basically creating a level of polish that has not been seen before." Later the same evening: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", "be revolutionary", and "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
The bar main set, and the bar this document holds every claim to: "impossible" is not available to any proof-of-work chain. The bar is that a majority of hash buys nothing: it cannot reverse what finality locked, cannot forge a proof the nodes re-execute, cannot change a rule without 95 percent signalling, and loses more than it earns. Every claim here is a number with a model or a simulation behind it, priced in rented hash at the measured 6 October rate (`docs/bench-log.md`, "Rental cost of hash, 6 October 2026": USD 0.0117 per MH/s-hour, so USD 11.7 per GH/s-hour and about USD 281 per GH/s-day; the live devnet at 1.16 GH/s), with its consequence per tier and what to build. Hours are agent hours (the project lead's rule: Claude-side work takes hours, never weeks).
The bar main set, and the bar this document holds every claim to: "impossible" is not available to any proof-of-work chain. The bar is that a majority of hash buys nothing: it cannot reverse what finality locked, cannot forge a proof the nodes re-execute, cannot change a rule without 95 percent signalling, and loses more than it earns. Every claim here is a number with a model or a simulation behind it, priced in rented hash at the measured 6 October rate (`docs/bench-log.md`, "Rental cost of hash, 6 October 2026": USD 0.0117 per MH/s-hour, so USD 11.7 per GH/s-hour and about USD 281 per GH/s-day; the live devnet at 1.16 GH/s), with its consequence per tier and what to build. Hours are agent hours (the founder's rule: Claude-side work takes hours, never weeks).
## The lanes
@ -17,11 +17,11 @@ The bar main set, and the bar this document holds every claim to: "impossible" i
| 7 frontier | `docs/analysis/horizon/frontier.md`; model `sim/horizon/frontier/frontier_model.py` | landed, commit 5ff7393 |
| 8 new-proof-of-work | `docs/analysis/horizon/new-pow.md`; prototypes `proto-newpow/` | landed, commits cbcff47 and 92db7c5 (two prototypes measured on rented 4090s, verdicts in section 6) |
## 1. One page for the project lead
## 1. One page for the founder
Closed 6 October 2026, 22:3x UK, every lane landed.
**Standing decisions (the project lead, 6 October 2026, 22:3x UK).** Verbatim: "Fees cannot fund security for a decade" and "Miners need to be the security". Meaning: miners are the security always, no time bound, no stake, no outside checkpoints or committee. The chain pays its own miners from emission plus fees; nobody pays upkeep, not the founder, not a treasury, not a dev fund. Self-sustaining means the emission curve keeps mining worth doing on its own for as long as fees are small, which lane 4 measures as a decade or more, so emission never decays on a schedule that assumes fees take over. Fee revenue is never assumed as the security budget in any model or public sentence. His third line, "Wrong constants and claims in our own text", acknowledges lane 6's finding; the nine ledger rows in item 9 are the fix.
**Standing decisions (the founder, 6 October 2026, 22:3x UK).** Verbatim: "Fees cannot fund security for a decade" and "Miners need to be the security". Meaning: miners are the security always, no time bound, no stake, no outside checkpoints or committee. The chain pays its own miners from emission plus fees; nobody pays upkeep, not the founder, not a treasury, not a dev fund. Self-sustaining means the emission curve keeps mining worth doing on its own for as long as fees are small, which lane 4 measures as a decade or more, so emission never decays on a schedule that assumes fees take over. Fee revenue is never assumed as the security budget in any model or public sentence. His third line, "Wrong constants and claims in our own text", acknowledges lane 6's finding; the nine ledger rows in item 9 are the fix.
1. **The proving pool was the one line where a majority earned more than it spent**: 11,636 IGN an hour at 51 percent of blocks. Fix: the proof verified in consensus. **In 0.3.16** (fork exec-sync-0313 da2d17ec), switch `proving_consensus_verify_daa` off by default. **Decision owed:** when it activates.
@ -41,11 +41,11 @@ Closed 6 October 2026, 22:3x UK, every lane landed.
9. **Polish.** Pause wording in Ember and the API: **in 0.3.16** (ember-tune b726ce4). `fork_is_close` and the publisher's activation guard: **in 0.3.16** (1357d28, 08f5276). Export-disk cap: **done** (gpu-fleet f9ad70e). Nine text corrections (M32, M33, F26, E19, E20, G15, P24, E21, P25) plus X31 to X33: **done**. Unsigned installers: **decision owed**. Relay fixes (28c028b) on fud-close: merge owed.
**Decisions owed from the project lead:** the cryptanalysis spend (USD 80,000 to 160,000); the testnet date word; the N ladder at genesis; the verification switch activation.
**Decisions owed from the founder:** the cryptanalysis spend (USD 80,000 to 160,000); the testnet date word; the N ladder at genesis; the verification switch activation.
## Decisions ready for the project lead (added 6 October 2026, 23:1x UK)
## Decisions ready for the founder (added 6 October 2026, 23:1x UK)
Two one-page verdicts landed after the close. Each is quoted verbatim from its file and each is **the project lead's decision owed**. The files sit on their branches and are not merged; read them there.
Two one-page verdicts landed after the close. Each is quoted verbatim from its file and each is **the founder's decision owed**. The files sit on their branches and are not merged; read them there.
### Emission: the tail
@ -188,7 +188,7 @@ The residual risks stated plainly: the first 20 days (no weight table yet); a pa
### Lane 6, polish
1. The pause had no cause on any surface: no `finality_reason` or frozen-table share in the node's report, the observer, the API, Ember or the hub (Q1, Q2, Q6); the 0.3.16 Ember carrier is on ember-tune.
2. Every update was urgent once an activation height was behind the node (`fork_is_close`, Q4), and nothing refused an activation height at or below the live DAA; both fixed on ember-tune for 0.3.16.
3. Unsigned installers on both desktops (Q3, the project lead's certificates) and the relay's security fixes still on fud-close (Q8).
3. Unsigned installers on both desktops (Q3, the founder's certificates) and the relay's security fixes still on fud-close (Q8).
### Lane 8, new-proof-of-work
1. Scheme A (mining is proving) is a bound, not a design: 2.9 MB of openings per block or a 32 to 40 ms verify against the 10 ms gate; the useful fraction is 8 percent at 1 GH/s and 0.08 percent at 100 GH/s.

View file

@ -305,11 +305,11 @@ Recommendation with numbers: hold the schedule as decided (2 GiB genesis, doubli
| 5a | A class v5 candidate `mx8 + sh256x35` with a shuffle-heavy shadow weight table (shfl 14, shfla 8 of 75), measured on the four owned cards before any cut: the first rung of proposal 5's ladder, plus the k-floor lever | the Mac's 5 percent point is 130,000; the shuffle mix raises the k floor 0.32 to 0.46 (approx) | section 5.3 | 4 to build the weight-table knob and packs, 1 Mac measure session (about 6 min under the lock, miner paused), 3 PC jobs | M5 Max -4.8 percent of rate at 37 W; 5090 -0.3 percent at its 431 W cap (a rig +23 percent electricity); 4070 0 at about 118 W; 9070 XT 0; verifier +0.23 ms Mac, +0.85 half-core | every owned card within 5 percent; bit-exact on three vendors; half-core verifier under 10 ms; the 5090's marginal pJ on the new mix read on three rungs |
| 6 | Hold the dataset schedule (2 GiB, years 4, 12, 28); write the prover footprint into the card-lifetime sentence | Steam shares and the measured prover peaks; one HBM3 stack holds every step | section 5.6 | 1 | 8 GB: mines to year 12, proves alone; 12 GB: mine-and-prove compressed to year 4, core-only to year 12, mines to year 28; 16 GB: compressed to year 4, core-only to year 12; 24 and 32 GB unconstrained to year 28 | the litepaper sentence matches the table; `docs/evidence.md` row "card lifetime" labelled designed |
| 7 | Make the Ember tune the shipped default per card model (the honest card's watts are the lever that moves every chip row) | the 4070 at 3.65 uJ untuned and 2.57 tuned (-30 percent); the 5090 2.65 bench against 2.34 app | section 5.1 | 2 (defaults table in the app from the fleet priors; already measured) | every NVIDIA tier gains 10 to 30 percent per joule; the chip's edge over the mid-tier falls from 8x to 13x toward 5x to 9x at v3 | MH per W per card model on the fleet night against the untuned baseline |
| 8 | Fund the k question: the item 3 cryptanalysis brief gains a chip-design line (a 14,000-lane SIMD array's energy per op on a random 32-lane program with shuffles, at N5 and at 28 nm) | every chip row at class v4 turns on k; nothing in the project measures it | section 5.3 | 0 agent hours; the project lead's money (part of the USD 80,000 to 160,000 brief) | none until the number lands; it decides whether 2x is reachable | a reviewed estimate of k with its range |
| 8 | Fund the k question: the item 3 cryptanalysis brief gains a chip-design line (a 14,000-lane SIMD array's energy per op on a random 32-lane program with shuffles, at N5 and at 28 nm) | every chip row at class v4 turns on k; nothing in the project measures it | section 5.3 | 0 agent hours; the founder's money (part of the USD 80,000 to 160,000 brief) | none until the number lands; it decides whether 2x is reachable | a reviewed estimate of k with its range |
Paragraphs.
1. The verifier gate is the one place tonight produced a measurement instead of a rule. The box proxy brackets a 2019 laptop from both sides (a 2022 server core at full boost; the same core with its sibling busy), and dr736 fails both brackets while class v4 passes both with 1.8 ms to spare. The measurement is one Windows build and one bench on the laptop the project lead already owns; until it lands, the half-core row replaces the "2.5x" from memory in every status file.
1. The verifier gate is the one place tonight produced a measurement instead of a rule. The box proxy brackets a 2019 laptop from both sides (a 2022 server core at full boost; the same core with its sibling busy), and dr736 fails both brackets while class v4 passes both with 1.8 ms to spare. The measurement is one Windows build and one bench on the laptop the founder already owns; until it lands, the half-core row replaces the "2.5x" from memory in every status file.
2. The FPGA lane's upper row was built on an activate rate (8 per 12 ns per channel) that the JEDEC HBM2 cycle table does not support (4 per 28 ns); the measured Shuhai rate sits exactly on the JEDEC ceiling. That reading can be wrong (the ICCAD table's clock interpretation, the half-bank count, the board watts are all approximate), which is why the F2 hour is the proposal and not the conclusion. It is cheap and it turns a public ceiling claim into a measured one.

View file

@ -176,7 +176,7 @@ Ways not in the seven, with the bound this lane gives:
| The 30-day window edges | (a) the frozen table expires at exactly day 30.00 after the last lock: both sides of a long split lock alone at once (M3); (b) a departed set leaves the sliding table over 30 days and the frozen one at the cliff (M4); (c) the first 30 days have no lock at all (3.8, `min_daa` = window); (d) new honest cohorts are under-weighted t/60 for 30 days (G) | stated in 3.7 items 2, 7, 9 | a partition or departure longer than one window ends with the fork of 3.7 item 9 and a manual F5 | | |
| The pause as a liveness attack | a silent set at or above 1/3 pauses every lock for as long as it stays silent (J, L1; sweep S) at zero marginal cost since it keeps earning | none in the rule; the node reports the pause; exchange guidance treats the chain as PoW with a 12-h depth | the 1/3 veto: 20 days at 51% | 0 once held | nothing directly; enables the 12-h PoW double spend below |
| What an attacker can do during a pause | plain proof of work: reorg up to the finality depth 43,200 DAA (12 h) with a heavier chain; beyond merge depth the honest blocks are abandoned (the 229-block shape); every certified checkpoint before the pause still binds | finality depth; the exchange guidance of 3.9 | 12 h of >50% hash | USD 146 / 1.5k / 15k / 146k | a deposit credited at the PoW depth; 559k IGN of subsidy as a miner |
| Tonight's departure (confirmed, lane 3 `finality-and-weight.md` 3.1 and 4.1) | 20 keys holding 42.7% of the frozen table stopped mining 18:27 to 18:30Z (the rehearsal job); the last lock 6842 at 18:39:40Z; 6843 determined with 53.1% of total signing and never locked; under v2 the stayers' sliding share crossed two thirds at 6912 (19:14:53Z, a 35-min pause) but Q5 held them at 57.3% of the frozen table; expected first lock when that table expires at DAA 216,402, about 20:40Z, or when 9.4 points of departed keys return. Sweep C agrees: 51% leaving pauses 10.5 days (v2) or 30.0 days (v3) at mainnet scale; at tonight's 46.9% (observer's view) 7.7 days under v2, 30 under v3 (lane 3, 4.1) | by design (F21: the project lead chose the pause over the fork); a view cannot tell a departure from a partition | anything over 1/3 of the table leaving at once pauses finality for a window | 0 | 0; what an attacker can do during it is the row above |
| Tonight's departure (confirmed, lane 3 `finality-and-weight.md` 3.1 and 4.1) | 20 keys holding 42.7% of the frozen table stopped mining 18:27 to 18:30Z (the rehearsal job); the last lock 6842 at 18:39:40Z; 6843 determined with 53.1% of total signing and never locked; under v2 the stayers' sliding share crossed two thirds at 6912 (19:14:53Z, a 35-min pause) but Q5 held them at 57.3% of the frozen table; expected first lock when that table expires at DAA 216,402, about 20:40Z, or when 9.4 points of departed keys return. Sweep C agrees: 51% leaving pauses 10.5 days (v2) or 30.0 days (v3) at mainnet scale; at tonight's 46.9% (observer's view) 7.7 days under v2, 30 under v3 (lane 3, 4.1) | by design (F21: the founder chose the pause over the fork); a view cannot tell a departure from a partition | anything over 1/3 of the table leaving at once pauses finality for a window | 0 | 0; what an attacker can do during it is the row above |
### 4.4 Miner signalling (P2)

View file

@ -6,7 +6,7 @@ What was read: `docs/spec/03-finality.md` (whole, 3.11 included), `04-seeds-and-
## 1. The three findings first
1. **Tonight's pause was the rule, not the aggregation path, and it was the frozen table that held it past 19:14Z.** The 20 keys that left the live chain between 17:20Z and 18:30Z held 3,026 of 7,083 blue blocks of the table frozen at the last lock (42.7 percent; the 13 that left with the 18:27 to 18:30Z rehearsal job alone 36.5 percent). The first unlocked checkpoint, 6843 at DAA 209,233 (about 18:40Z), had 75 of 93 voters' votes and 53.1 percent of total weight on the observer's node, under the two-thirds floor; certificates had kept forming for 18 checkpoints while node 1 and the observer were down (6824 to 6842, 18:30 to 18:39Z, 83 signers, 78.5 to 79.7 percent of total). From 19:14:53Z (checkpoint 6912) the stayers held 74.9 percent of the sliding table and still did not lock, because they hold 57.3 percent of the frozen table of lock 6842, which stands until DAA 216,402 (about 20:40Z). Under rule v2 the first lock would have come at 6912, 35 minutes after the last; under v3 the pause is one window, 2 hours on the devnet and 30 days on mainnet (spec 3.7 item 2, the price the project lead took on 4 October).
1. **Tonight's pause was the rule, not the aggregation path, and it was the frozen table that held it past 19:14Z.** The 20 keys that left the live chain between 17:20Z and 18:30Z held 3,026 of 7,083 blue blocks of the table frozen at the last lock (42.7 percent; the 13 that left with the 18:27 to 18:30Z rehearsal job alone 36.5 percent). The first unlocked checkpoint, 6843 at DAA 209,233 (about 18:40Z), had 75 of 93 voters' votes and 53.1 percent of total weight on the observer's node, under the two-thirds floor; certificates had kept forming for 18 checkpoints while node 1 and the observer were down (6824 to 6842, 18:30 to 18:39Z, 83 signers, 78.5 to 79.7 percent of total). From 19:14:53Z (checkpoint 6912) the stayers held 74.9 percent of the sliding table and still did not lock, because they hold 57.3 percent of the frozen table of lock 6842, which stands until DAA 216,402 (about 20:40Z). Under rule v2 the first lock would have come at 6912, 35 minutes after the last; under v3 the pause is one window, 2 hours on the devnet and 30 days on mainnet (spec 3.7 item 2, the price the founder took on 4 October).
2. **Of the four candidate rules, only the departure announcement keeps the one-third bound.** In the simulator (3 seeds, mainnet scale) the decaying denominator and the hysteresis floor both restore liveness after tonight's departure in under an hour and both reopen the partition double lock (fast decay: both sides of every 360-minute partition lock alone from minute 120, 467 to 473 conflicting locks in the 50/50 honest split and 17 to 264 in the poisoned eclipse; slow decay: both sides of the 12-day splits lock alone at day 0.5 to 0.9, 31,545 to 32,246 conflicts; hysteresis: a 20 percent equivocator conflicts from minute 60, 565 to 597 locks, the 13.3 percent bound of 3 October back). The leave rule locks 1 hour after the departure (0.04 days; under 4 devnet minutes) with 0 conflicting locks in every partition, eclipse and equivocator row, and an attacker who buys keys to make them leave gains nothing it would not get by signing with them (w + L must still reach 2/3). The two-tier report never conflicts in its final tier by construction and shows 516 to 1,062 conflicting PROVISIONAL locks in every 360-minute partition and about 34,000 in the 12-day splits, so it is a reporting layer with a health warning, not a rule.
3. **Weight costs USD 8,424 x N x W / (1 - W) to rent for the full window** at the measured USD 11.7 per GH/s-hour: a veto (34 percent) against a 1 GH/s network is USD 4,300 over 30 days (0.52 x N of hash, a +52 percent step on the chart from day 1), against 1 TH/s USD 4.3 M; locking alone (67 percent) is 2.03 x N for 30 days (USD 17,100 per GH/s of network, USD 17 M at 1 TH/s), and faster is dearer (22 days: 10.6 x N). Buying old keys costs the seller's own rental equivalent, decays to nothing in 30 days (sim K), and nothing in the protocol makes weight unbuyable; what keeps the price at the rental cost is that the seller keeps a copy and one equivocation strips the key.

View file

@ -1,6 +1,6 @@
# Horizon lane 7: frontier. Predictions to 2030, what no proof-of-work chain has shipped, and what Igneum can
6 October 2026, evening UK. Lane 7 of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon`, at origin/master 3f4f719). the project lead's words: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", and "be revolutionary". Main's bar: not features, but ideas that change what a proof-of-work chain is or what a GPU owner is to the world, each with its evidence, cost, gate, and the attack a Monero or Kaspa core developer would mount. I argue each attack as the project's own four personas would hear it (`.claude/agents/cryptographer.md`, `consensus-engineer.md`, `execution-engineer.md`, `miner-community-lead.md`).
6 October 2026, evening UK. Lane 7 of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon`, at origin/master 3f4f719). The founder's words: "research anything else that we can research too, predictions, forward thinking, what we can actually do that has not been done or applied, think outside the box", and "be revolutionary". Main's bar: not features, but ideas that change what a proof-of-work chain is or what a GPU owner is to the world, each with its evidence, cost, gate, and the attack a Monero or Kaspa core developer would mount. I argue each attack as the project's own four personas would hear it (`.claude/agents/cryptographer.md`, `consensus-engineer.md`, `execution-engineer.md`, `miner-community-lead.md`).
Nothing in this file is a prediction of the coin's price, an offer to sell anything, or a change to any consensus parameter. Every chip figure is arithmetic on cited memory and logic figures; every GPU figure names its bench entry; a figure from memory says approximate. No em dashes.
@ -12,7 +12,7 @@ Nothing in this file is a prediction of the coin's price, an offer to sell anyth
## 0. Everything ranked by payoff over difficulty
Payoff 1 to 5 is what the idea does for the chain's security, the coin's utility or the GPU owner's position in the world, if it works. Difficulty is Claude-side hours to a measurable prototype (the project lead's rule: hours, never weeks). Verdicts: do now, prototype, watch, never. The "never" rows carry a sharp reason so the rest are not fantasy.
Payoff 1 to 5 is what the idea does for the chain's security, the coin's utility or the GPU owner's position in the world, if it works. Difficulty is Claude-side hours to a measurable prototype (the founder's rule: hours, never weeks). Verdicts: do now, prototype, watch, never. The "never" rows carry a sharp reason so the rest are not fantasy.
| Rank | Idea | Payoff | Hours | Verdict | One line why |
|---|---|---|---|---|---|
@ -322,7 +322,7 @@ The at-risk amount is five to nine orders of magnitude above the designed coin b
### 3.6 Treasury-less audit funding: bounties from the burn, a review escrow on upgrades
**The idea.** the project lead removed the dev fund (spec 5.5) and the project pays audits from the Ember dev fee and founders' mined coins (litepaper). The question: money for audits that comes from users paying for something, with no standing address. Two mechanisms. (a) **Burn redirect.** The base fee burns to nobody. A reproducible break submitted under spec 0.5 and accepted by 60 percent of blue blocks over a window redirects the base-fee burn of the next 7 (execution) or 30 (consensus) days to the submitter's address, once, then returns to burning. No address exists between events. (b) **Review escrow.** An upgrade proposal under 5.7 must escrow IGN in a contract that pays reviewers named in the proposal on a 60 percent "review complete" signal, or refunds on failure; the proposer pays, which is a user paying for a thing (the right to propose code).
**The idea.** the founder removed the dev fund (spec 5.5) and the project pays audits from the Ember dev fee and founders' mined coins (litepaper). The question: money for audits that comes from users paying for something, with no standing address. Two mechanisms. (a) **Burn redirect.** The base fee burns to nobody. A reproducible break submitted under spec 0.5 and accepted by 60 percent of blue blocks over a window redirects the base-fee burn of the next 7 (execution) or 30 (consensus) days to the submitter's address, once, then returns to burning. No address exists between events. (b) **Review escrow.** An upgrade proposal under 5.7 must escrow IGN in a contract that pays reviewers named in the proposal on a 60 percent "review complete" signal, or refunds on failure; the proposer pays, which is a user paying for a thing (the right to propose code).
**Why nobody shipped it.** Zcash funds development from the block subsidy (NU6: 8 percent to Zcash Community Grants, 12 percent to a protocol lockbox, ZIP 1015); Decred from a 10 percent treasury spent by stakeholder vote, capped at 4 percent of balance a month since January 2026 (DCP-0013); Monero from the CCS, donations off-chain; Optimism from an 850 M OP reserve for retro funding. Bug bounties pay 10 percent of funds at risk (Immunefi's standard) from the protocol's own treasury; Code4rena runs contests at zero platform fee since 2025. Nobody funds audits from a burn redirect, because a burn redirect is a subsidy to a payee by rule, and the chains that wanted that built a treasury. New in form; a dev fund in substance (see the Monero attack).
@ -341,7 +341,7 @@ At launch traffic a 30-day redirect is under one audit contest; at half-full blo
**Hours.** 20 for the escrow contract and the redirect rule as a proposal kind; 0 for the honest alternative, which already exists.
**The gate.** None that a simulator settles; the gate is the project lead's: does a per-event, miner-approved payee with no standing address pass the test that removed the dev fund?
**The gate.** None that a simulator settles; the gate is the founder's: does a per-event, miner-approved payee with no standing address pass the test that removed the dev fund?
**Per tier.** Miners vote on each payout with their blocks and can refuse all of them; holders see supply that would have burned paid to a named person; a prover, rig, pool user and rollup customer see nothing unless a break affects them; the node operator gains a proposal kind.
@ -415,7 +415,7 @@ At launch traffic a 30-day redirect is under one audit contest; at half-full blo
**The idea as asked.** Make the leader election depend in part on proving work, so the energy that picks the block maker is useful.
**Why every attempt failed, cited.** Primecoin (2013) found Cunningham and bi-twin prime chains that nobody uses (Bitcoin Magazine, July 2013). Gridcoin pays for BOINC work and stops if BOINC stops (gridcoin.us; the 2022 "Challenges of PoUW" survey, arXiv 2209.03865). Ball, Rosen, Sabin and Vasudevan (eprint 2017/203) gave proofs of useful work from fine-grained problems (Orthogonal Vectors, 3SUM, APSP) and state the conditions: the problem must be sampleable at a tunable hardness with instances the miner cannot choose, and the verifier must be cheaper than the work. Ofelimos (Fitzi, Kiayias, Panagiotakos, Russell, CRYPTO 2022, eprint 2021/1379) got a provably secure protocol by making the work a doubly efficient local search whose usefulness is a side effect and small. The 2026 "Economics of Proof-of-Useful-Work" (arXiv 2606.06700) and the empirical study of Pearl's cuPOW (arXiv 2606.04819, "The Usefulness Gap") find the same gap between the work paid for and the work anyone wanted. I found no Coinbase paper on the subject (searched 6 October 2026); if the project lead has one in mind, its title is needed. Aleo ran proving as the consensus work and the fastest prover won (CLAUDE.md: the Aleo lesson; litepaper precedents table, approximate). Boundless's PoVW (docs.boundless.network/zkc/mining/overview) pays ZKC pro rata to cycles proven per epoch with a stake that scales with the work, which is a reward for proving, not a leader election, and it is on a proof-of-stake chain.
**Why every attempt failed, cited.** Primecoin (2013) found Cunningham and bi-twin prime chains that nobody uses (Bitcoin Magazine, July 2013). Gridcoin pays for BOINC work and stops if BOINC stops (gridcoin.us; the 2022 "Challenges of PoUW" survey, arXiv 2209.03865). Ball, Rosen, Sabin and Vasudevan (eprint 2017/203) gave proofs of useful work from fine-grained problems (Orthogonal Vectors, 3SUM, APSP) and state the conditions: the problem must be sampleable at a tunable hardness with instances the miner cannot choose, and the verifier must be cheaper than the work. Ofelimos (Fitzi, Kiayias, Panagiotakos, Russell, CRYPTO 2022, eprint 2021/1379) got a provably secure protocol by making the work a doubly efficient local search whose usefulness is a side effect and small. The 2026 "Economics of Proof-of-Useful-Work" (arXiv 2606.06700) and the empirical study of Pearl's cuPOW (arXiv 2606.04819, "The Usefulness Gap") find the same gap between the work paid for and the work anyone wanted. I found no Coinbase paper on the subject (searched 6 October 2026); if the founder has one in mind, its title is needed. Aleo ran proving as the consensus work and the fastest prover won (CLAUDE.md: the Aleo lesson; litepaper precedents table, approximate). Boundless's PoVW (docs.boundless.network/zkc/mining/overview) pays ZKC pro rata to cycles proven per epoch with a stake that scales with the work, which is a reward for proving, not a leader election, and it is on a proof-of-stake chain.
**The sampleability problem, plainly.** A lottery needs a puzzle whose instances are drawn at random from a distribution the miner cannot steer, whose hardness is tunable by a target, and whose solution is verifiable in milliseconds. zkVM proving has none of these: the instances (segments, jobs) are chosen by users and producers, the hardness is whatever the program is, and the verifier is tens of milliseconds to seconds. Any blend ("a miner's lottery target eases in proportion to its proven cycles last hour") gives the fastest prover more blocks, which is Aleo with a cap, and a cap small enough to be safe is a reward too small to be useful.
@ -640,10 +640,10 @@ Smaller than the sections above; each with hours and a gate.
| I6 | A finality-pause page on the site that shows the connected weight fraction live, so the "node reports the pause" sentence has a public face | 4 | Shows tonight's 18:42Z pause from the observer's data | Tonight's incident |
| I7 | Equivocation-evidence bounty paid in sortition slots: the key that first carries valid evidence inherits the stripped key's shard assignments for 30 days (no coins move; weight is reassigned, not created) | 12 | Two signers under one key on the fast-time harness; the evidence carrier wins the stripped key's draws | Makes watching for equivocation pay without a treasury |
| I8 | Mandatory proofs activation height set from a measured coverage share (spec 7.8 item 10) | 4 | Coverage above 99 percent for 7 days on the devnet | The rule is written and off |
| I9 | The exclusive window at 25 s and the claim timeout at 120 s on the phase 4 devnet (decided by the project lead, P9) with the economy simulator re-run at the measured shard times from `prover-tiers-real-cards.md` instead of the 20-s target | 6 | The 3060 class's shard share within 5 points of its weight share | The inputs changed today |
| I9 | The exclusive window at 25 s and the claim timeout at 120 s on the phase 4 devnet (decided by the founder, P9) with the economy simulator re-run at the measured shard times from `prover-tiers-real-cards.md` instead of the 20-s target | 6 | The 3060 class's shard share within 5 points of its weight share | The inputs changed today |
| I10 | `eth_getProof`, `debug_traceTransaction`, `eth_subscribe` (D5 step 2) before any outside team | 24 | Foundry's debugger and the Blockscout fork run against a devnet node | The light client and every tool depend on `eth_getProof` |
| I11 | Register chain ids 4461 to 4463 on ethereum-lists/chains before the public testnet (spec 7.1) | 1 | The PR merged | Wallets |
| I12 | Publish the 2028 tier table (section 2.6) on the miner page with its three rates, so no card owner buys on a promise | 2 | Live | the project lead's consequences rule |
| I12 | Publish the 2028 tier table (section 2.6) on the miner page with its three rates, so no card owner buys on a promise | 2 | Live | the founder's consequences rule |
| I13 | A spec sentence in 03 and 05: "no coin stake; the only thing at stake is 30 days of public work" (4.1) | 1 | Text | Before 3.2 is prototyped |
| I14 | The litepaper's income table gains the proving-market arithmetic of 3.11 in one line | 1 | Text | Ledger P6 asked for honesty; the number makes it concrete |
| I15 | A ledger entry beside E4 recording 3.6 as considered and rejected on E4's ground | 1 | Text | So the question is not re-asked |

View file

@ -6,7 +6,7 @@ What was read: `docs/spec/02-consensus.md` (2.1 parameters, 2.3 the difficulty r
## 1. The question and the answer in one paragraph
the project lead asked for Kaspa's answer to solo-miner variance: a higher block rate. Run A ran Devnet 2 at 10 blocks per second through one seed and produced 77 percent red blocks, 321 tips and a 55-block reorg. The propagation model says the links and the star did not do that: with the measured latencies it predicts under 0.1 percent red at 10 bps in a star and in a mesh. What did it is the seed's CPU per block, measured at 61 ms (narrow DAG) to 345 ms (mergeset 150 to 200), against a budget of 100 ms per block at 10 bps; with that cost in the model the star gives 47 to 88 percent red and queueing waits of 26 to 1,769 s, which are the "Accepted 100 blocks via relay" batches in the log. The difficulty rule then read blue work over a chain step capped at 2 s and hardened until the DAG ran at 2 s / (chain-step spacing) of target (model 4 blocks/s at a 5-s spacing, record 3.3 to 3.6), while a narrower-but-still-wide DAG would have made it ease (the direction main reported); counting every mergeset block over the real span, as Kaspa's window does, is unbiased in both regimes. The block rate for the public testnet is 1 bps; 10 bps is a gated step that needs the per-block node cost under 50 ms on a laptop core at a mergeset of 248, the checkpoint interval and the clock cap re-denominated in DAA seconds, and vote aggregation, because with C1 in blue blocks 8,192 voters at 10 bps are 66 GB per node per day of votes.
The founder asked for Kaspa's answer to solo-miner variance: a higher block rate. Run A ran Devnet 2 at 10 blocks per second through one seed and produced 77 percent red blocks, 321 tips and a 55-block reorg. The propagation model says the links and the star did not do that: with the measured latencies it predicts under 0.1 percent red at 10 bps in a star and in a mesh. What did it is the seed's CPU per block, measured at 61 ms (narrow DAG) to 345 ms (mergeset 150 to 200), against a budget of 100 ms per block at 10 bps; with that cost in the model the star gives 47 to 88 percent red and queueing waits of 26 to 1,769 s, which are the "Accepted 100 blocks via relay" batches in the log. The difficulty rule then read blue work over a chain step capped at 2 s and hardened until the DAG ran at 2 s / (chain-step spacing) of target (model 4 blocks/s at a 5-s spacing, record 3.3 to 3.6), while a narrower-but-still-wide DAG would have made it ease (the direction main reported); counting every mergeset block over the real span, as Kaspa's window does, is unbiased in both regimes. The block rate for the public testnet is 1 bps; 10 bps is a gated step that needs the per-block node cost under 50 ms on a laptop core at a mergeset of 248, the checkpoint interval and the clock cap re-denominated in DAA seconds, and vote aggregation, because with C1 in blue blocks 8,192 voters at 10 bps are 66 GB per node per day of votes.
## 2. Method

View file

@ -2,7 +2,7 @@
6 October 2026, evening UK, lane `new-proof-of-work`, worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon` from master, at 3f4f719). Output of this lane: this file and `proto-newpow/<scheme>/`. Nothing here touches the shipped hash, `igneum-pow`, the node, the manifest or the live devnet; every prototype is a benchmark beside the worker, never inside it.
the project lead's mandate, verbatim: "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
The founder's mandate, verbatim: "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
## 0. Progress (kept current for the coordinator)
@ -189,7 +189,7 @@ Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and e
What the rows say:
1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5).
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised) [corrected 8 October 2026: these per-MAC figures count 1,024 multiply-adds per tile per lane where a tile is 1,024 per warp and 32 per lane, so they are low by 32x: 22 pJ at R = 8 falling to 1.8 pJ at R = 512; the watts and microjoules per hash stand; counter-asic-4-research.md 15.1a at 71fd465b], 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force.
4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
@ -273,7 +273,7 @@ Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `
| Scheme | Verdict | Why, in one line |
|---|---|---|
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add as first counted, 1.8 to 22 pJ with the per-lane count corrected on 8 October 2026, 15 W for 4,096 tiles per hash) is the reason, and the correction strengthens it: a chip's MAC at 0.04 to 0.4 pJ (claimed) against the GPU's 1.8 pJ gives a chip k of 0.03 to 0.3 on tile work, below the ALU shadow's. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
@ -286,7 +286,7 @@ Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `
## 7. Ranked next steps
Hours are agent hours (the project lead's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own. Ranks 2, 3 and 4 were written as B's gates before the ladder landed; after section 5 they are WITHDRAWN (B is not carried forward as class content; the rows stay so the reasoning is visible) and the live order is 1, 5, 6, 7, 8.
Hours are agent hours (the founder's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own. Ranks 2, 3 and 4 were written as B's gates before the ladder landed; after section 5 they are WITHDRAWN (B is not carried forward as class content; the rows stay so the reasoning is visible) and the live order is 1, 5, 6, 7, 8.
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|---|---|---|---|---|---|---|
@ -314,11 +314,11 @@ One paragraph each.
| Question | Why it could not be closed tonight | What closes it |
|---|---|---|
| The 2019-class core (O-1.14) | no such core in the fleet; igneum-build-1 is Zen 4, the Mac is M5 Max; the 2.5x rule stands in | rank 7 |
| The AMD WMMA fragment layout | PC 1 is the project lead's desk and the AMD rows were owed all day (status file); no AMD card on RunPod or Vast tonight (fleet agent) | rank 2 |
| The AMD WMMA fragment layout | PC 1 is the founder's desk and the AMD rows were owed all day (status file); no AMD card on RunPod or Vast tonight (fleet agent) | rank 2 |
| The Apple emulation of mm8 | a Metal emulation kernel is a 3-hour job and the Mac measure lock was free; not started because the AMD gate decides first whether B proceeds | rank 3 |
| The canonical state serialisation for C | a design item that touches the exec layer (`igneum/exec`), out of this lane's files | rank 1 |
| The pruning-proof witness for C's historical PoW | spec 02 and 10 items | rank 1 |
| Whether the tensor path's marginal energy on the 5090 differs from the 4090's | one card measured (box 1); the 5090 is on the project lead's desk | a PC 2 job with the same `run.sh` |
| Whether the tensor path's marginal energy on the 5090 differs from the 4090's | one card measured (box 1); the 5090 is on the founder's desk | a PC 2 job with the same `run.sh` |
| The Ampere row (3060, 3080, 3090) | every Ampere card of the fleet was mining and proving the live devnet; a loaded 3090 was offered and declined (a loaded card's rate is not a number) | one quiet Ampere pod |
| The verifier with a byte-dot instruction (VNNI, NEON udot) | the C reference is scalar | 1 h: an AVX-VNNI and a NEON path in `verify_ref.c`, measured on both cores |
| Scheme C's leaf array on an 8 GB card at the 2 GiB design size | the build holds dataset + cache + leaves on the device (4.3 GiB) unless chunked | the chunked build (section 5.2 says whether it is trivial) |

View file

@ -2,7 +2,7 @@
Date: 6 October 2026, evening UK (written 20:00 to 21:00Z, while the live devnet's finality was still paused). Lane 6 of the Horizon programme. Worktree `/Users/joshm/Projects/igneum-wt-horizon` (branch `horizon`, HEAD c3aa502). Output: this file only. No file outside it was edited; nothing was built, deployed, posted or started; the installed app and every live service were read, never touched.
the project lead's bar: "a level of polish that has not been seen before." This file is the honest audit: what each shipped system does, the named comparator and the exact screen or feature it has that we lack or do worse, the ledger rows that close the gap (95 rows, ids Q1 to Q105 with gaps, each with a file or screen, a severity, hours of agent work, an owner and a gate), and what is already better than the comparator. The first ten rows are what the project lead will notice first when he opens each product tomorrow.
The founder's bar: "a level of polish that has not been seen before." This file is the honest audit: what each shipped system does, the named comparator and the exact screen or feature it has that we lack or do worse, the ledger rows that close the gap (95 rows, ids Q1 to Q105 with gaps, each with a file or screen, a severity, hours of agent work, an owner and a gate), and what is already better than the comparator. The first ten rows are what the founder will notice first when he opens each product tomorrow.
## 0. What was read and run
@ -14,26 +14,26 @@ Run (nothing live, nothing built): `node --test app/igneum-app/ui/*.test.mjs` (3
Comparator claims are from memory unless a repository or page is named, and are labelled approximate.
## 1. The ten rows the project lead will notice first
## 1. The ten rows the founder will notice first
| Rank | Id | Row | File or screen | Severity | Hours | Owner | Who it hits | Gate |
|---|---|---|---|---|---|---|---|---|
| 1 | Q1 | The pause has no cause anywhere. The node reports `finality_reason` (`active`, `window filling, N of M`, `paused`) but not the frozen-table share that held tonight's pause; the observer drops even the reason; the API has no reason field; `/live` computes "N% of weight silent" from the sliding table, which read 11 percent silent at 20:05Z while finality was paused (held by the table frozen at lock 6842, 57.3 percent signing). Add `finality_reason`, `held_by` (frozen index, its signing share, its expiry DAA) and lane 3's `finality_provisional` end to end | `vendor/igneum-node-0310/consensus/src/processes/finality.rs:1794-1807` (reason exists, no frozen share), `tools/observer/observer.mjs:720` (copies `finalityActive` only), `site/api/live.mjs:114-122` (no reason), `site/live.html:520` (the silent percentage) | the project lead-visible | 4 (node 1.5, observer and API 1, `/live` and `/api/stats` copy 1, spec 3.9 row 0.5) | consensus-engineer (node), site owner (observer, API, page) | holder, exchange: the one line that says whether a pause is a silent third, a split, or a table waiting to expire; miner: nothing lost, but the app can finally explain itself; operator: a pause with a cause and an end time | A forced pause on Devnet 2 (one slice over a third stopped): `/api/live.finality` carries `reason`, `held_by`, `provisional`; `/live` reads "finality paused since 18:40Z: 89% of the sliding weight is signing, held by the table frozen at lock 6842 (57% signing) until DAA 216,402 (about 20:40Z)"; the exchange guidance in spec 3.9 carries the provisional row (lane 3 proposal 4) |
| 2 | Q2 | Ember never says "paused". The engine keeps only `last_lock`, `last_lock_at`, `age_s`, `votes`; the node line showed "#6842 · 1 h ago" under "a point the miners agreed can never be undone" for the whole two-hour pause. No `finality_active` anywhere in the app | `app/igneum-app/src/state.rs:247-248`, `src/engine.rs:3495-3497, 3896-3902` (lock parsed from the miner's `LOCK checkpoint` line only), `ui/app.js:1350-1351` | the project lead-visible | 3 | app owner (engine reads `getFinalityCheckpoints` or the miner's `FINALITY` line, spec 3.9 row; UI state and view test) | home miner and rig: the app tells them finality is paused and why, instead of a lock age that climbs; pool user: the same on the pool page later | Mock state with `finality_active:false, reason:"paused"` renders "Finality paused since 18:40Z, 89% of the 30-day weight signing, held by the frozen table until about 20:40Z" on the node line and in the pill; `view.test.mjs` case; seen on the Devnet 2 forced pause |
| 3 | Q3 | A fresh machine's first 60 seconds start with a warning. macOS: ad hoc signature, no Developer ID, no notarisation; the engine strips `com.apple.quarantine` from its own bundle on start; the README's "Right-click > Open" bypass is gone on macOS 15 (approximate). Windows: installer and exes unsigned, SmartScreen "Windows protected your PC", then a UAC prompt for the firewall rule 20 to 50 s in before any screen explains it | `packaging/mac/build-dmg.sh:118-122`, `packaging/mac/README.md:51-53`, `app/igneum-app/src/main.rs:78`, `packaging/windows/README.md:76, 111`, `packaging/windows/build-installer.ps1:188`, `app/igneum-app/src/ota.rs:981-1015` | the project lead-visible | 6 (Developer ID signing plus `notarytool` and stapling in `build-dmg.sh` 3; Authenticode `signtool` step in `windows.yml` and `build-installer.ps1` 2; firewall step after `setup_done` with a sentence on the Cards screen 1), plus the project lead: an Apple Developer account and an OV or EV code-signing certificate (purchases, hours of the project lead's time, approximate) | app owner; the project lead (the certificates) | every new miner on every tier, Windows and macOS: Signal's and Tailscale's installers open with no warning (approximate); today ours open with two | Fresh macOS 15 VM: the DMG's app opens on double-click, `spctl -a -vv` says accepted and notarised; fresh Windows 11 VM: no SmartScreen interstitial, `signtool verify /pa` passes; the first UAC prompt appears after the Cards screen names it |
| 4 | Q4 | Every update is "urgent" once an activation height is behind the node: `fork_is_close(Some(198000), 209000)` is true, so the manifest's stale `activation 198000` turned the 0.3.14 update into "the node is 0 blocks away. Installing now", stripped Later and skipped every safe-moment guard (the PC 1 install under a measurement job at 17:52:54Z). The same path has no finality input: an update applies through a pause | `app/igneum-app/src/manifest.rs:350-355` (`daa + 1800 >= h`), `:322-347` (`safe_to_apply`, no finality), `src/ota.rs:484`, `ui/app.js:217`, `docs/plans/release-0.3.14.md:79`, `release-0.3.15.md` section 6 | the project lead-visible | 3 (close means within 1,800 blocks ahead and not behind 1; the publisher drops a passed `activation_height` 0.5; `Moment.finality_paused` holds a non-urgent update 1.5) | app owner | home miner, rig: no surprise restart mid-measurement or mid-pause; operator: an activation that has passed is not an emergency | Unit tests: behind by any amount is not close; `publish-manifest.sh` refuses a passed height; a paused mock holds the update with the words "finality is paused; installing when it resumes"; the 0.3.16 manifest carries no stale height |
| 5 | Q5 | The homepage says "final" while finality is paused, in a 90-word hero, under a share card that renders as a thumbnail. The chain scene takes `src.locked` and draws the dashed line labelled "final" at the newest locked block with no `finality.active` check (R4.6.3 was fixed on `/live` only); the hero is one 90-word paragraph with two bench deep links and "ships when the packaging row lands"; the miner section still says "a 24 GB NVIDIA card proves as well" while the hero and litepaper say 8 GB; home, litepaper, live and explorer share a 256 px `og-small.png` with `summary` cards | `site/index.html:810, 838` (final label), `:345` (hero), `:494` and `site/miner.html:7, 13, 21, 265, 382` (24 GB), `site/index.html:15-19` (OG) | the project lead-visible | 3 (final label 0.5, hero 1, 24 GB sentence 0.5, four 1200x630 cards 1) | site owner | every visitor; a holder or exchange reading "final" during a pause is the worst of them | No "final" label while `finality.active` is false (the `/live` rule); hero under 40 words with one link; one proving-tier sentence on every page; every page's shared card is 1200x630 |
| 6 | Q6 | The hub shows stale data as live, said "active" for the first 20 unlocked checkpoints, says "paused" with no since or cause, and buried the pause under 401 miner_quiet and miner_back events. After the first load an API failure only changes the header to "offline: ..."; tiles, cards and "Refreshed" keep the old values. `finality_active` flips only after `presence_window` (20 on the devnet) indices without a lock, so 18:42 to 18:50Z read "active" in white with no lock forming. No `finality_paused` or `finality_resumed` event exists; `live_events` between 18:25Z and 21:30Z holds 203 `miner_quiet`, 198 `miner_back`, 55 `difficulty`, 30 `checkpoint_locked` (the last at 18:42:10Z) and no pause row. The checkpoint table also showed "% of active" above 100 (102.4 to 113.4 percent on indices 6900 to 6930) | `relay/ui.html:564` (stale), `:367-368` (tile), `relay/api/console.mjs:214`, `vendor/igneum-node-0310/consensus/src/processes/finality.rs:1794` and `vendor/igneum-node/consensus/core/src/finality.rs:63, 75` (`presence_window` 240 mainnet, 20 devnet), `tools/observer/observer.mjs:652` (the only finality event), `live_checkpoints.fraction_active` rows 6900 to 6930 | the project lead-visible | 4 (dim every panel with "last data N min ago" 1; amber on the first under-2/3 checkpoint 0.5; `finality_paused` and `finality_resumed` events, and collapse quiet/back pairs under 5 min 1.5; clamp or explain the active fraction 1) | fleet agent (hub), consensus-engineer (observer events, the fraction) | operator: the hub is the project lead's first screen; a holder reading the public API gets the same events | Cut the network: every hub panel dims within 15 s; a forced Devnet 2 pause turns the tile amber on the first checkpoint under two thirds, posts one `finality_paused` event with the cause and one `finality_resumed` with the duration; no fraction above 100 percent on any row for 24 h |
| 7 | Q7 | Discord said nothing. The bot has no credentials file and its timer is not installed, so tonight's pause produced zero posts; had it been live, a condition already firing at the watcher's first look is marked `preexisting` and never opens an incident; the pulse hides the pause in its description with no colour or title change | `docs/community/discord-hooks.md:81-82`, `tools/community/discord-hooks.mjs:481-485`, `:252-255`, `:158`, `infra/build-server/discord-hooks/install.sh` | the project lead-visible | 3 (install and one live pulse 1; a pre-existing condition opens with "since at least <first look>" 1; paused pulse in a distinct colour with a title suffix 1) | miner-community-lead (community owner) | every Discord reader, which on launch day is every miner | `check` prints three "set"; one live pulse in #numbers; a fixture where the pause predates the first tick opens an incident; the paused pulse renders amber with "(finality paused)" in the title |
| 1 | Q1 | The pause has no cause anywhere. The node reports `finality_reason` (`active`, `window filling, N of M`, `paused`) but not the frozen-table share that held tonight's pause; the observer drops even the reason; the API has no reason field; `/live` computes "N% of weight silent" from the sliding table, which read 11 percent silent at 20:05Z while finality was paused (held by the table frozen at lock 6842, 57.3 percent signing). Add `finality_reason`, `held_by` (frozen index, its signing share, its expiry DAA) and lane 3's `finality_provisional` end to end | `vendor/igneum-node-0310/consensus/src/processes/finality.rs:1794-1807` (reason exists, no frozen share), `tools/observer/observer.mjs:720` (copies `finalityActive` only), `site/api/live.mjs:114-122` (no reason), `site/live.html:520` (the silent percentage) | the founder-visible | 4 (node 1.5, observer and API 1, `/live` and `/api/stats` copy 1, spec 3.9 row 0.5) | consensus-engineer (node), site owner (observer, API, page) | holder, exchange: the one line that says whether a pause is a silent third, a split, or a table waiting to expire; miner: nothing lost, but the app can finally explain itself; operator: a pause with a cause and an end time | A forced pause on Devnet 2 (one slice over a third stopped): `/api/live.finality` carries `reason`, `held_by`, `provisional`; `/live` reads "finality paused since 18:40Z: 89% of the sliding weight is signing, held by the table frozen at lock 6842 (57% signing) until DAA 216,402 (about 20:40Z)"; the exchange guidance in spec 3.9 carries the provisional row (lane 3 proposal 4) |
| 2 | Q2 | Ember never says "paused". The engine keeps only `last_lock`, `last_lock_at`, `age_s`, `votes`; the node line showed "#6842 · 1 h ago" under "a point the miners agreed can never be undone" for the whole two-hour pause. No `finality_active` anywhere in the app | `app/igneum-app/src/state.rs:247-248`, `src/engine.rs:3495-3497, 3896-3902` (lock parsed from the miner's `LOCK checkpoint` line only), `ui/app.js:1350-1351` | the founder-visible | 3 | app owner (engine reads `getFinalityCheckpoints` or the miner's `FINALITY` line, spec 3.9 row; UI state and view test) | home miner and rig: the app tells them finality is paused and why, instead of a lock age that climbs; pool user: the same on the pool page later | Mock state with `finality_active:false, reason:"paused"` renders "Finality paused since 18:40Z, 89% of the 30-day weight signing, held by the frozen table until about 20:40Z" on the node line and in the pill; `view.test.mjs` case; seen on the Devnet 2 forced pause |
| 3 | Q3 | A fresh machine's first 60 seconds start with a warning. macOS: ad hoc signature, no Developer ID, no notarisation; the engine strips `com.apple.quarantine` from its own bundle on start; the README's "Right-click > Open" bypass is gone on macOS 15 (approximate). Windows: installer and exes unsigned, SmartScreen "Windows protected your PC", then a UAC prompt for the firewall rule 20 to 50 s in before any screen explains it | `packaging/mac/build-dmg.sh:118-122`, `packaging/mac/README.md:51-53`, `app/igneum-app/src/main.rs:78`, `packaging/windows/README.md:76, 111`, `packaging/windows/build-installer.ps1:188`, `app/igneum-app/src/ota.rs:981-1015` | the founder-visible | 6 (Developer ID signing plus `notarytool` and stapling in `build-dmg.sh` 3; Authenticode `signtool` step in `windows.yml` and `build-installer.ps1` 2; firewall step after `setup_done` with a sentence on the Cards screen 1), plus the founder: an Apple Developer account and an OV or EV code-signing certificate (purchases, hours of the founder's time, approximate) | app owner; the founder (the certificates) | every new miner on every tier, Windows and macOS: Signal's and Tailscale's installers open with no warning (approximate); today ours open with two | Fresh macOS 15 VM: the DMG's app opens on double-click, `spctl -a -vv` says accepted and notarised; fresh Windows 11 VM: no SmartScreen interstitial, `signtool verify /pa` passes; the first UAC prompt appears after the Cards screen names it |
| 4 | Q4 | Every update is "urgent" once an activation height is behind the node: `fork_is_close(Some(198000), 209000)` is true, so the manifest's stale `activation 198000` turned the 0.3.14 update into "the node is 0 blocks away. Installing now", stripped Later and skipped every safe-moment guard (the PC 1 install under a measurement job at 17:52:54Z). The same path has no finality input: an update applies through a pause | `app/igneum-app/src/manifest.rs:350-355` (`daa + 1800 >= h`), `:322-347` (`safe_to_apply`, no finality), `src/ota.rs:484`, `ui/app.js:217`, `docs/plans/release-0.3.14.md:79`, `release-0.3.15.md` section 6 | the founder-visible | 3 (close means within 1,800 blocks ahead and not behind 1; the publisher drops a passed `activation_height` 0.5; `Moment.finality_paused` holds a non-urgent update 1.5) | app owner | home miner, rig: no surprise restart mid-measurement or mid-pause; operator: an activation that has passed is not an emergency | Unit tests: behind by any amount is not close; `publish-manifest.sh` refuses a passed height; a paused mock holds the update with the words "finality is paused; installing when it resumes"; the 0.3.16 manifest carries no stale height |
| 5 | Q5 | The homepage says "final" while finality is paused, in a 90-word hero, under a share card that renders as a thumbnail. The chain scene takes `src.locked` and draws the dashed line labelled "final" at the newest locked block with no `finality.active` check (R4.6.3 was fixed on `/live` only); the hero is one 90-word paragraph with two bench deep links and "ships when the packaging row lands"; the miner section still says "a 24 GB NVIDIA card proves as well" while the hero and litepaper say 8 GB; home, litepaper, live and explorer share a 256 px `og-small.png` with `summary` cards | `site/index.html:810, 838` (final label), `:345` (hero), `:494` and `site/miner.html:7, 13, 21, 265, 382` (24 GB), `site/index.html:15-19` (OG) | the founder-visible | 3 (final label 0.5, hero 1, 24 GB sentence 0.5, four 1200x630 cards 1) | site owner | every visitor; a holder or exchange reading "final" during a pause is the worst of them | No "final" label while `finality.active` is false (the `/live` rule); hero under 40 words with one link; one proving-tier sentence on every page; every page's shared card is 1200x630 |
| 6 | Q6 | The hub shows stale data as live, said "active" for the first 20 unlocked checkpoints, says "paused" with no since or cause, and buried the pause under 401 miner_quiet and miner_back events. After the first load an API failure only changes the header to "offline: ..."; tiles, cards and "Refreshed" keep the old values. `finality_active` flips only after `presence_window` (20 on the devnet) indices without a lock, so 18:42 to 18:50Z read "active" in white with no lock forming. No `finality_paused` or `finality_resumed` event exists; `live_events` between 18:25Z and 21:30Z holds 203 `miner_quiet`, 198 `miner_back`, 55 `difficulty`, 30 `checkpoint_locked` (the last at 18:42:10Z) and no pause row. The checkpoint table also showed "% of active" above 100 (102.4 to 113.4 percent on indices 6900 to 6930) | `relay/ui.html:564` (stale), `:367-368` (tile), `relay/api/console.mjs:214`, `vendor/igneum-node-0310/consensus/src/processes/finality.rs:1794` and `vendor/igneum-node/consensus/core/src/finality.rs:63, 75` (`presence_window` 240 mainnet, 20 devnet), `tools/observer/observer.mjs:652` (the only finality event), `live_checkpoints.fraction_active` rows 6900 to 6930 | the founder-visible | 4 (dim every panel with "last data N min ago" 1; amber on the first under-2/3 checkpoint 0.5; `finality_paused` and `finality_resumed` events, and collapse quiet/back pairs under 5 min 1.5; clamp or explain the active fraction 1) | fleet agent (hub), consensus-engineer (observer events, the fraction) | operator: the hub is the founder's first screen; a holder reading the public API gets the same events | Cut the network: every hub panel dims within 15 s; a forced Devnet 2 pause turns the tile amber on the first checkpoint under two thirds, posts one `finality_paused` event with the cause and one `finality_resumed` with the duration; no fraction above 100 percent on any row for 24 h |
| 7 | Q7 | Discord said nothing. The bot has no credentials file and its timer is not installed, so tonight's pause produced zero posts; had it been live, a condition already firing at the watcher's first look is marked `preexisting` and never opens an incident; the pulse hides the pause in its description with no colour or title change | `docs/community/discord-hooks.md:81-82`, `tools/community/discord-hooks.mjs:481-485`, `:252-255`, `:158`, `infra/build-server/discord-hooks/install.sh` | the founder-visible | 3 (install and one live pulse 1; a pre-existing condition opens with "since at least <first look>" 1; paused pulse in a distinct colour with a title suffix 1) | miner-community-lead (community owner) | every Discord reader, which on launch day is every miner | `check` prints three "set"; one live pulse in #numbers; a fixture where the pause predates the first tick opens an incident; the paused pulse renders amber with "(finality paused)" in the title |
| 8 | Q8 | The relay's security fixes are not on master. Commit 28c028b (X23 to X28: `relay/lib/guard.mjs`, `handler.mjs`, run signatures, retention) is on `fud-close` and the `ledger-*` branches only; `git merge-base --is-ancestor 28c028b master` is false. On this tree any of the three intake keys (one ships in every miner app) can post, sync and delete console items, read the whole feed, and every client puts the token in the URL path | `relay/api/console.mjs:249, 283-301`, `relay/lib/relay.mjs:34-46`, `relay/clients/agent.sh:10`, `igneum-agent.ps1:13, 72, 154`, `tools/relay.mjs:24` | operator-visible (security) | 2 (merge or cherry-pick with the 47 relay tests 1; handler test that an intake key gets 403 on post, sync, delete 1) | fleet agent (relay owner) | operator: the console is the control plane for every PC and the fleet; a miner app's intake key is in every install | `28c028b` is an ancestor of the deploying branch; 47 relay tests pass; the handler test above passes; `relay/README.md` names the deployed commit |
| 9 | Q86 | The fleet page is blind to the standing fleet, to death and to finality. `page.py:43-47` publishes `standing` and `devnet2` and the page references neither (grep 0); a dead box reads "running" for up to 20 minutes because state comes from `boxes.json` and `standing.jsonl` is never read; no fleet or console script reads `finality_active`, so tonight's pause was found by hand. The night the project lead ruled that 22 boxes stay up permanently, the page that shows them cannot show them | `igneum-wt-gpu-fleet/tools/fleet/page.py:38-47`, dlsite `index.html:243`, `lib/standing.py:54-56`, `tools/console.mjs:124` | the project lead-visible | 6 (standing block with USD per day and the Devnet 2 gate line 2; "unreachable since HH:MM" within 2 minutes 2; finality on the page and in `console.mjs chain` 2) | fleet agent | operator (the project lead reads this page every evening); every tier indirectly: a dead standing box is weight that left silently, the class of tonight's pause | the page shows the standing count and spend; a box killed by hand reads "unreachable since" within 2 minutes; a forced Devnet 2 pause reads on the page and in the console |
| 10 | Q9 | The wallet page describes 0.1.4 while UI 3 (five state words, light mode, pounds line) sits unmerged on `wallet-ui-3`; the MetaMask guide's testnet RPC is still labelled a placeholder, its devnet RPC port is the miner's node (26790) while a wallet-only Mac runs 26800, and `wallet_addEthereumChain` carries `blockExplorerUrls: []` and no `iconUrls`, so MetaMask shows no explorer link and a blank icon | `site/wallet.html:261-266, 312`, `igneum-wt-wallet-ui/docs/plans/wallet-ui-3.md:93, 112`, `site/metamask.html:197, 225, 233, 285-286`, `igneum-wt-wallet/app/igneum-wallet/src/node.rs:24-26` | the project lead-visible | 6 (wallet 0.1.5 with UI 3 published and the page rewritten 4; guide: both ports, the placeholder line removed, explorer URL and icon in the request 2) | app owner (wallet), site owner (guide) | holder: the first non-miner product; a wrong port or a placeholder RPC is a dead end at the first step | `downloads.json` wallet-mac 0.1.5; the page names pending, included, executed, proven, finalised; `eth_chainId` returns 0x116e from the guide's URL; MetaMask shows the explorer link after the add-chain button |
| 9 | Q86 | The fleet page is blind to the standing fleet, to death and to finality. `page.py:43-47` publishes `standing` and `devnet2` and the page references neither (grep 0); a dead box reads "running" for up to 20 minutes because state comes from `boxes.json` and `standing.jsonl` is never read; no fleet or console script reads `finality_active`, so tonight's pause was found by hand. The night the founder ruled that 22 boxes stay up permanently, the page that shows them cannot show them | `igneum-wt-gpu-fleet/tools/fleet/page.py:38-47`, dlsite `index.html:243`, `lib/standing.py:54-56`, `tools/console.mjs:124` | the founder-visible | 6 (standing block with USD per day and the Devnet 2 gate line 2; "unreachable since HH:MM" within 2 minutes 2; finality on the page and in `console.mjs chain` 2) | fleet agent | operator (the founder reads this page every evening); every tier indirectly: a dead standing box is weight that left silently, the class of tonight's pause | the page shows the standing count and spend; a box killed by hand reads "unreachable since" within 2 minutes; a forced Devnet 2 pause reads on the page and in the console |
| 10 | Q9 | The wallet page describes 0.1.4 while UI 3 (five state words, light mode, pounds line) sits unmerged on `wallet-ui-3`; the MetaMask guide's testnet RPC is still labelled a placeholder, its devnet RPC port is the miner's node (26790) while a wallet-only Mac runs 26800, and `wallet_addEthereumChain` carries `blockExplorerUrls: []` and no `iconUrls`, so MetaMask shows no explorer link and a blank icon | `site/wallet.html:261-266, 312`, `igneum-wt-wallet-ui/docs/plans/wallet-ui-3.md:93, 112`, `site/metamask.html:197, 225, 233, 285-286`, `igneum-wt-wallet/app/igneum-wallet/src/node.rs:24-26` | the founder-visible | 6 (wallet 0.1.5 with UI 3 published and the page rewritten 4; guide: both ports, the placeholder line removed, explorer URL and icon in the request 2) | app owner (wallet), site owner (guide) | holder: the first non-miner product; a wrong port or a placeholder RPC is a dead end at the first step | `downloads.json` wallet-mac 0.1.5; the page names pending, included, executed, proven, finalised; `eth_chainId` returns 0x116e from the guide's URL; MetaMask shows the explorer link after the add-chain button |
Next in line, outside the ten: Q10 (the explorer knows nothing about finality and swallows failures after first load, 3 hours), Q80 (the Windows installer pipeline is dead on GitHub billing) and Q12 (Ember's light mode fails contrast on the private key and every live state).
Hours for the ten: 40 agent hours, plus the project lead's certificate purchases for Q3.
Hours for the ten: 40 agent hours, plus the founder's certificate purchases for Q3.
## 2. Tonight's pause on every surface (the ledger row the project lead asked for)
## 2. Tonight's pause on every surface (the ledger row the founder asked for)
Times UTC. The chain-side facts are lane 3's (`finality-and-weight.md` 3.1) and the observer's own rows, read tonight: last lock 6842 at about 18:39:40Z; 6843 proposed at 53.1 percent of total; `finality_active` false from about 18:50Z (index 6862, twenty indices without a lock, `presence_window` 20); no lock after 6842 as of 20:05:03Z (`live_state.updated_at`), `total_weight` 7,167, `active_weight` 6,346 (88.5 percent signing on the sliding table), `voters` 85, and no `finality_reason` key in the stored JSON. Sampled `live_checkpoints` rows:
@ -82,20 +82,20 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q2 | Finality pause invisible (section 1) | `state.rs:247`, `engine.rs:3495`, `app.js:1350` | the project lead-visible | 3 | app owner | section 1 |
| Q4 | Every update urgent once the activation is behind; no finality hold (section 1) | `manifest.rs:322-355` | the project lead-visible | 3 | app owner | section 1 |
| Q2 | Finality pause invisible (section 1) | `state.rs:247`, `engine.rs:3495`, `app.js:1350` | the founder-visible | 3 | app owner | section 1 |
| Q4 | Every update urgent once the activation is behind; no finality hold (section 1) | `manifest.rs:322-355` | the founder-visible | 3 | app owner | section 1 |
| Q11 | Offline start says "The update failed." A check error with no held manifest calls `set_error`; the cached manifest is loaded for the floor but never into `self.manifest`; the notice offers "Try again", which POSTs install | `ota.rs:649-654, 195-200`, `app.js:29, 1198` | user-visible | 1 | app owner | notices test: a check-class error renders "Could not check for updates" with no install action |
| Q12 | Light mode fails contrast on every live text: molten `#B8731F` on white 3.81:1, on bone 3.38:1; ember `#E04A14` on bone 3.61:1; molten carries the private key, the address and the live states. `miner-ui-3.md` section 9 claims 4.5:1 everywhere | `app.css:29, 263, 338, 431` | user-visible | 1 | app owner | a node script over the token pairs asserts 4.5:1 for every text token in both schemes; CI runs it |
| Q13 | Per-card rejects and faults never reach the row (the lolMiner line) | `app.js:291-309, 1290` | the project lead-visible on PC 1 | 2 | miner-community-lead | row meta "N rejected · M faults" when nonzero; view test |
| Q13 | Per-card rejects and faults never reach the row (the lolMiner line) | `app.js:291-309, 1290` | the founder-visible on PC 1 | 2 | miner-community-lead | row meta "N rejected · M faults" when nonzero; view test |
| Q14 | Windows first run raises a UAC prompt (firewall rule) before any screen explains it (part of Q3) | `ota.rs:981-1015`, `index.html:71-99` | user-visible | 1 | app owner | the firewall step runs after `setup_done`; the Cards screen names it |
| Q15 | Mac first-run instruction stale: "Right-click > Open" (part of Q3) | `packaging/mac/README.md:51`, `site/miner.html` download copy | the project lead-visible on a fresh Mac | 1 | app owner | the download page and README name the Privacy and Security step until the certificate lands |
| Q15 | Mac first-run instruction stale: "Right-click > Open" (part of Q3) | `packaging/mac/README.md:51`, `site/miner.html` download copy | the founder-visible on a fresh Mac | 1 | app owner | the download page and README name the Privacy and Security step until the certificate lands |
| Q16 | No rate history: a 10-minute strip against HiveOS's hours | `app.js:1052-1127` | user-visible | 4 | app owner | one-hour per-card sparkline from an engine ring buffer; view test |
| Q17 | Hill climb hidden until `tune_climb` is on the state | `index.html:341`, `app.js:1407` | operator-visible | 1 | app owner | the row shows whenever the engine answers `tune/goal` |
| Q18 | Earnings never shows mined IGN; "£0.00 earned" is the first number on the tab | `index.html:250`, `app.js:1364` | user-visible | 2 | miner-community-lead | the tab shows blocks mined and the subsidy they earned in IGN, with the devnet line under it |
| Q19 | 10 px type in five places, 9 px ruler | `app.css:196, 271, 272, 302, 491, 531` | cosmetic | 1 | app owner | no `font-size` under 11 px except the ruler |
| Q20 | Design screens in `docs/design/app-screens/*.png` are a UI 1 app (v0.3.0 tiles, a Finality card the shipped UI lacks) | `docs/design/app-screens/` | cosmetic | 1 | app owner | screens regenerated from `?screen=` on the current build |
| Q21 | AMD step line prints a percent as watts ("owed: a unit word for AMD") | `ember-tune.md:278` | operator-visible | 1 | miner-community-lead | the line reads "70% (an offset)"; unit test |
| Q22 | A code comment naming the founder ships in the UI bundle (`// The list is ordered by performance (the project lead, 6 October 2026)`), and the ui bundle is outside every forbidden-string check | `app/igneum-app/ui/app.js:270`, `tools/ci/identity-check.sh` | cosmetic (identity rule) | 0.5 | app owner | `tools/ci` forbidden-string check covers `app/igneum-app/ui/`; the comment reads "(ruling of 6 October 2026)" |
| Q22 | A code comment naming the founder ships in the UI bundle (`// The list is ordered by performance (the founder, 6 October 2026)`), and the ui bundle is outside every forbidden-string check | `app/igneum-app/ui/app.js:270`, `tools/ci/identity-check.sh` | cosmetic (identity rule) | 0.5 | app owner | `tools/ci` forbidden-string check covers `app/igneum-app/ui/`; the comment reads "(ruling of 6 October 2026)" |
| Q23 | Two copy-law slips: "Votes lock the chain; leave it on." (aphorism), "The window can close. The miner keeps going" (two-beat) | `index.html:404, 81` | cosmetic | 0.5 | app owner | reworded; the grep below stays at 0 em dashes |
**Error states, exact strings.** Engine down (four failed polls): pill "engine away" (`app.js:1437`); node line "Node: no answer", sub "the engine is not answering; the window reconnects by itself"; Mac host "The engine stopped. Quit and open Igneum Miner again." (`IgneumMiner.swift:233`); Windows "The engine stopped. Close this window and open Igneum Miner again." (`host.cpp:301`). Node down: "the node could not start" or the engine message, "the node is not running", "the external node went away" (`app.js:452-453`, `engine.rs:3163`); pill "node failed" (`app.js:1428`); digest box "the node is not running" (`app.js:1341`); canvas "waiting for the node to sync" (`app.js:1099`). Finality paused: nothing (Q2). Update server unreachable: "The update failed. Manifest signature: <curl error>." (`app.js:1360`, Q11). GPU lost: "Card removed: <name>. Its worker stopped." (`app.js:95`), row word "removed", sub "unplugged; its worker stopped. The row goes in five minutes." (`app.js:303`); Code 43: "<name>: not usable (Code 43). No worker runs on it." with "reboot with the card attached; if it persists, reinstall the driver with the card attached" (`app.js:100`, `detect.rs:55`).
@ -123,10 +123,10 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q9 | Page vs UI 3; guide placeholders; empty explorer URL (section 1) | `site/wallet.html`, `site/metamask.html` | the project lead-visible | 6 | app owner, site owner | section 1 |
| Q9 | Page vs UI 3; guide placeholders; empty explorer URL (section 1) | `site/wallet.html`, `site/metamask.html` | the founder-visible | 6 | app owner, site owner | section 1 |
| Q30 | Private key export copies to the clipboard with no clear and no timer | `igneum-wt-wallet/.../ui/app.js:129, 463` | user-visible | 2 | app owner | clipboard cleared after 60 s by the host; unit test on the timer |
| Q31 | Wallet OTA trusts one key; no `revoked_keys`, unlike the miner's updater | `docs/plans/consequences-2026-10-05.md:36`, `updater.rs:7, 729` | operator-visible | 3 | app owner | updater carries `OTA_PUBLIC_KEYS [K1, K2]` and honours `revoked_keys`; the three-key test of the rig installer |
| Q32 | No hardware wallet path despite spec 8.5 "MUST offer" | `docs/spec/08-client-security.md:36` | the project lead-visible | 24, or 0.5 to relabel | app owner (cryptographer reviews) | a Ledger signs one devnet transfer through the Ethereum app at chain id 4463; or the spec row reads "Designed, not shipped" |
| Q32 | No hardware wallet path despite spec 8.5 "MUST offer" | `docs/spec/08-client-security.md:36` | the founder-visible | 24, or 0.5 to relabel | app owner (cryptographer reviews) | a Ledger signs one devnet transfer through the Ethereum app at chain id 4463; or the spec row reads "Designed, not shipped" |
| Q33 | Single account, no second address, no address book | `engine.rs:343`, `ui/index.html:177` | user-visible | 6 | app owner | two derived accounts, a saved recipient reused in Send |
| Q34 | `confirm()` on Remove wallet never renders in the Mac host | `app.js:598`, `wallet-ui-3-audit.md:189` | user-visible | 1 | app owner | in-page two-step card; snapshot test |
| Q35 | Finality pause: node card "paused" or "not checkpoint-checkable without a node", rows stay "in block N"; public-RPC mode never says what to do | wallet `app.js:185, 321, 340` | user-visible | 2 | app owner | the row reads "in block N; finality is paused on the network" and the public-RPC line names the fix with a button |
@ -157,11 +157,11 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q6 | Stale shown as live; "active" for 20 unlocked indices; "paused" with no since or cause; event flood (section 1) | `relay/ui.html:564, 367-368` | the project lead-visible | 4 | fleet agent, consensus-engineer | section 1 |
| Q6 | Stale shown as live; "active" for 20 unlocked indices; "paused" with no since or cause; event flood (section 1) | `relay/ui.html:564, 367-368` | the founder-visible | 4 | fleet agent, consensus-engineer | section 1 |
| Q8 | X23 to X28 fixes off master; intake key can write and delete; token in URL (section 1) | `relay/api/console.mjs:249`, `relay/lib/relay.mjs:44` | operator-visible | 2 | fleet agent | section 1 |
| Q40 | A hung job reads "running" forever: no elapsed time, no timeout, no "last line N min ago" | `relay/api/console.mjs:182-187`, `relay/ui.html:317` | the project lead-visible | 2 | fleet agent | a killed job reads "no report for 12 min" in red within one refresh |
| Q40 | A hung job reads "running" forever: no elapsed time, no timeout, no "last line N min ago" | `relay/api/console.mjs:182-187`, `relay/ui.html:317` | the founder-visible | 2 | fleet agent | a killed job reads "no report for 12 min" in red within one refresh |
| Q41 | Jobs tab trusts the unsigned `igneum-jobs.json` while the apps verify the signed file | `relay/api/console.mjs:170` vs `tools/jobs.mjs:57-60` | operator-visible | 1.5 | fleet agent | the tab shows the envelope's signature state per file |
| Q42 | Job result modal: no error lines first, no duration, no anchors (the Actions view) | `relay/ui.html:319`, `console.mjs:184-187` | the project lead-visible | 3 | fleet agent | a failed run opens with its five error lines and its duration first |
| Q42 | Job result modal: no error lines first, no duration, no anchors (the Actions view) | `relay/ui.html:319`, `console.mjs:184-187` | the founder-visible | 3 | fleet agent | a failed run opens with its five error lines and its duration first |
| Q43 | Machines: no per-machine key state, no revoke, roles never enforced (the Tailscale view) | `relay/lib/relay.mjs:14`, `relay/api/relay.mjs:191-195` | operator-visible | 3 (after Q8) | fleet agent | a revoked PC's register is 403 and its card says "key revoked" |
| Q44 | Builds: no rollback, no manifest history (the Vercel view) | `relay/ui.html:328-341` | operator-visible | 4 | fleet agent | one tap republishes the previous manifest and the fleet's OTA state shows it |
| Q45 | "live feed unreachable: live feed unreachable" doubled; the whole Chain tab empties, losing Events and the infra card | `relay/api/console.mjs:210`, `relay/ui.html:357` | cosmetic | 0.5 | fleet agent | live feed down: tiles grey, Events and infra card still render |
@ -200,7 +200,7 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q5 | "final" during a pause; 90-word hero; 24 GB contradiction; thumbnail OG (section 1) | `index.html:345, 494, 810, 838, 15-19` | the project lead-visible | 3 | site owner | section 1 |
| Q5 | "final" during a pause; 90-word hero; 24 GB contradiction; thumbnail OG (section 1) | `index.html:345, 494, 810, 838, 15-19` | the founder-visible | 3 | site owner | section 1 |
| Q50 | No downloads page: hashes only for HiveOS, no signing key, no verify line, the Linux button goes to an anchor | `index.html:363-365, 516`, `miner.html:549-558` | user-visible | 2 | site owner | `/download` lists every artefact with version, size, sha256, the OTA key fingerprint and a verify command per OS; the Linux button links the file |
| Q51 | `/miners` has 6 rows, two card models, "not measured" MH/W on every row; the eleven-card fleet table is not ingested | `miners.html:183`, `miner-bench.json` | user-visible | 2 | miner-community-lead | rows carry W and MH/W for every measured card |
| Q52 | `dl\.igneum` in `forbidden-strings.txt` matches the public download host on three built pages; the scrub guards only bench and miners | `forbidden-strings.txt:25`, `build.mjs:171, 401` | operator-visible | 0.5 | site owner | the pattern becomes `dl\.igneum\.network/dl/(?!public/)` or `/dl/[0-9a-f]{12,}`, and the check runs on every built page |
@ -210,10 +210,10 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Q56 | Light-section contrast fails (3.08:1 links, 3.97:1 eyebrow) | `index.html:119-120, 130` | cosmetic | 0.5 | site owner | AA on every token pair; the same node script as Q12 |
| Q57 | "devnet v0" eyebrow on the live page; the chain is v4 | `live.html:219` | cosmetic | 0.1 | site owner | matches the app's network word |
| Q58 | Heading order (h3 before the first h2) | `index.html:359` | cosmetic | 0.2 | site owner | h2 or a styled div |
| Q59 | No team, custody or key-powers page (Reddit round 4 artefact 8); "0 admin keys in consensus" tile unchanged | `index.html:611`; no file | user-visible | 1 (plus the project lead's policy call) | site owner | the key-powers table published; the tile links it |
| Q60 | Roadmap dates versus the testnet: the journey says "Public testnet, Aug to Oct 2027" and the litepaper "Pools and the public testnet are August 2027", while `docs/plans/testnet-go.md` has `igneum-testnet-1` seeds up and a go checklist dated 5 October 2026 | `index.html:679`, `litepaper.html:719`, `docs/plans/testnet-go.md:1-4` | the project lead-visible | 0.5 (after the project lead decides which is true) | site owner | one date for the public testnet on the journey, the litepaper and the download notice |
| Q59 | No team, custody or key-powers page (Reddit round 4 artefact 8); "0 admin keys in consensus" tile unchanged | `index.html:611`; no file | user-visible | 1 (plus the founder's policy call) | site owner | the key-powers table published; the tile links it |
| Q60 | Roadmap dates versus the testnet: the journey says "Public testnet, Aug to Oct 2027" and the litepaper "Pools and the public testnet are August 2027", while `docs/plans/testnet-go.md` has `igneum-testnet-1` seeds up and a go checklist dated 5 October 2026 | `index.html:679`, `litepaper.html:719`, `docs/plans/testnet-go.md:1-4` | the founder-visible | 0.5 (after the founder decides which is true) | site owner | one date for the public testnet on the journey, the litepaper and the download notice |
| Q61 | No Content-Security-Policy header on the site (the relay has none either); inline scripts throughout, two pinned CDN modules on the homepage | `site/vercel.json:7`, `relay/vercel.json:31-43`, `verify/verify.js:5-6` | operator-visible (security) | 2 | site owner | a CSP with hashes or nonces for the inline scripts and `script-src` limited to self and cdn.jsdelivr; every page renders with no console violation |
| Q62 | Copy-law borderline headlines: "Mined by GPUs. Proven by fire." (the brand line, the project lead's call), "GPUs are back · for good", "Last hour's chip is already obsolete.", "Your coins. Final means final.", "Dates slip. Gates do not.", "Install. Start. The card mines and proves.", "Nothing is mined here." | `index.html:343, 344, 424`, `wallet.html` h1, `litepaper.html:678`, `miner.html` h1, `404.html` h1 | cosmetic | 1 | site owner (the project lead rules on the brand line) | each either kept by the project lead's word or reworded |
| Q62 | Copy-law borderline headlines: "Mined by GPUs. Proven by fire." (the brand line, the founder's call), "GPUs are back · for good", "Last hour's chip is already obsolete.", "Your coins. Final means final.", "Dates slip. Gates do not.", "Install. Start. The card mines and proves.", "Nothing is mined here." | `index.html:343, 344, 424`, `wallet.html` h1, `litepaper.html:678`, `miner.html` h1, `404.html` h1 | cosmetic | 1 | site owner (the founder rules on the brand line) | each either kept by the founder's word or reworded |
**Reddit round 4, still open in the HTML.** Finding 6 (24 GB): half fixed (Q5). Finding 7 (eleven-card table): open (Q51). Finding 8 (admin keys): sentence added to the litepaper (`:659`), tile unchanged, no table (Q59). Finding 15 ("Monero's idea, finished for GPUs"): open (`index.html:444`, `litepaper.html:429`). Finding 16 ("devnet v0"): open (Q57). Finding 22: `/ledger` present, no nav item. Findings 2, 3, 5, 9, 10, 12, 13, 17, 19: fixed in the current files. "0% anyone else in the protocol" legend still live (`index.html:604`).
@ -266,8 +266,8 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q67 | Member transport is plaintext TCP against spec 9.3's TLS 1.3 with session binding; `binding` ignored | `server.rs:1-2, 115`, `protocol.rs:58-60` | user-visible (hijack risk) | 8 | consensus-engineer | rustls listener; a replayed `authorize` on a second connection is refused (O-9.7) |
| Q68 | Node down shows "difficulty 0, DAA 0" and a ramp-day-0 reward; no "node unreachable" state | `node.rs:350-368`, `web/index.html:163-164` | the project lead-visible (once deployed) | 2 | miner-community-lead | the page renders "Node unreachable since <time>" against a stopped node |
| Q69 | Raw integers (DAA, block count, shares) with no separators; difficulty uses `toLocaleString()` with no locale, so it varies by browser; hash rate 2 decimals against the site's 1 | `web/index.html:152, 163, 172, 190` | the project lead-visible | 1 | miner-community-lead | every number formatted as section 4.3 says |
| Q68 | Node down shows "difficulty 0, DAA 0" and a ramp-day-0 reward; no "node unreachable" state | `node.rs:350-368`, `web/index.html:163-164` | the founder-visible (once deployed) | 2 | miner-community-lead | the page renders "Node unreachable since <time>" against a stopped node |
| Q69 | Raw integers (DAA, block count, shares) with no separators; difficulty uses `toLocaleString()` with no locale, so it varies by browser; hash rate 2 decimals against the site's 1 | `web/index.html:152, 163, 172, 190` | the founder-visible | 1 | miner-community-lead | every number formatted as section 4.3 says |
| Q70 | Ledger is a JSON file rewritten every 15 s; hash-rate samples and check cost not persisted, so a restart zeroes every rate tile and the 24-hour luck | `state.rs:704-710`, `main.rs:131-144` | user-visible | 4 | miner-community-lead | restart test: tiles recover within one sample window; README names the loss window |
| Q71 | Address lookup never shows the payment history the API returns | `api.rs:148`, `web/index.html:181-194` | user-visible | 1 | miner-community-lead | payments table under the lookup |
| Q72 | No per-miner history, no charts, no worker-offline notice, no `/metrics`; `/health` always ok | `state.rs:779-781`, `web/index.html:187`, `main.rs:152-174`, `api.rs:195` | user-visible, operator-visible | 4 + 4 + 2 + 1 | miner-community-lead | hourly buckets per address kept 7 days with a sparkline; opt-in webhook after 10 min silence; Prometheus text endpoint; health reflects `net.synced` and template age |
@ -292,7 +292,7 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q7 | Bot not live; pre-existing never opens; pulse hides the pause (section 1) | `discord-hooks.mjs:481-485, 252-255, 158` | the project lead-visible | 3 | miner-community-lead | section 1 |
| Q7 | Bot not live; pre-existing never opens; pulse hides the pause (section 1) | `discord-hooks.mjs:481-485, 252-255, 158` | the founder-visible | 3 | miner-community-lead | section 1 |
| Q74 | No block-production stall condition, the actual first signal of tonight's 18:30Z gap | `:437-472`, `fixtures/discord/live.json` | user-visible | 2 | miner-community-lead | `blocks_stalled` condition, 3 minutes, fixture test |
| Q75 | Resolve posts a new message instead of editing the open one | `:706-712` | cosmetic | 2 | miner-community-lead | PATCH `/messages/<id>` with the resolve fields |
| Q76 | Embeds link only `/live` and `/miner`; the last lock could link `/block/<hash>` | `:271, 297, 350` | cosmetic | 1 | miner-community-lead | last lock links its block page |
@ -318,9 +318,9 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q86 | The fleet page is blind to the standing fleet, to death and to finality. `page.py:43-47` publishes `standing` and `devnet2`, and the page's `index.html` references neither (grep 0; it reads boxes, results, phases, night, log, `:243`); a dead box reads "running" for up to 20 minutes because state comes from `boxes.json` (`page.py:38`) and `standing.jsonl` is never read; no fleet or console script reads `finality_active` (only `bps-collect.py:30` tails "Finality:" log lines), so tonight's pause was found by hand (`CLAUDE.md`, the standing-fleet rule) | `page.py:38-47`, dlsite `index.html:243`, `lib/standing.py`, `tools/console.mjs:124` | the project lead-visible | 6 (standing block, USD per day, Devnet 2 height and last gate line 2; "unreachable since HH:MM" from `standing.jsonl` within 2 minutes 2; `getFinalityCheckpoints` in the standing check and "finality paused since, reason" on the page and the console 2) | fleet agent | the page shows the standing count and spend; a box killed by hand reads "unreachable since" within 2 minutes; a forced Devnet 2 pause reads on the page and in `console.mjs chain` |
| Q87 | `rerent` does not set up or supervise the new box: the docstring promises both (`lib/standing.py:14-16`), the code rents and patches rows (`:78-91`); a re-rented card bills idle | `lib/standing.py:78-91` | the project lead-visible (spend) | 2 | fleet agent | the new label shows synced and mining on the page within the install time |
| Q88 | One-shot boxes have no autostop; `cap_vast 1000` and `cap_runpod 500` enforce nothing (`page.py:45`; `vast.py:61-72` has no cap check) | `page.py:45`, `vast.py:61-72`, `autorun.py` | the project lead-visible (spend) | 1.5 | fleet agent | `autorun` destroys a one-shot box N minutes after its done line; `rent` refuses above the cap |
| Q86 | The fleet page is blind to the standing fleet, to death and to finality. `page.py:43-47` publishes `standing` and `devnet2`, and the page's `index.html` references neither (grep 0; it reads boxes, results, phases, night, log, `:243`); a dead box reads "running" for up to 20 minutes because state comes from `boxes.json` (`page.py:38`) and `standing.jsonl` is never read; no fleet or console script reads `finality_active` (only `bps-collect.py:30` tails "Finality:" log lines), so tonight's pause was found by hand (`CLAUDE.md`, the standing-fleet rule) | `page.py:38-47`, dlsite `index.html:243`, `lib/standing.py`, `tools/console.mjs:124` | the founder-visible | 6 (standing block, USD per day, Devnet 2 height and last gate line 2; "unreachable since HH:MM" from `standing.jsonl` within 2 minutes 2; `getFinalityCheckpoints` in the standing check and "finality paused since, reason" on the page and the console 2) | fleet agent | the page shows the standing count and spend; a box killed by hand reads "unreachable since" within 2 minutes; a forced Devnet 2 pause reads on the page and in `console.mjs chain` |
| Q87 | `rerent` does not set up or supervise the new box: the docstring promises both (`lib/standing.py:14-16`), the code rents and patches rows (`:78-91`); a re-rented card bills idle | `lib/standing.py:78-91` | the founder-visible (spend) | 2 | fleet agent | the new label shows synced and mining on the page within the install time |
| Q88 | One-shot boxes have no autostop; `cap_vast 1000` and `cap_runpod 500` enforce nothing (`page.py:45`; `vast.py:61-72` has no cap check) | `page.py:45`, `vast.py:61-72`, `autorun.py` | the founder-visible (spend) | 1.5 | fleet agent | `autorun` destroys a one-shot box N minutes after its done line; `rent` refuses above the cap |
| Q89 | Supervisor restart loop without backoff or a failure line | `box-standing.sh:46-49` | operator-visible | 1.5 | fleet agent | after N failures a `standing_node_failed` line and a page flag |
| Q90 | Dropped miner and GPU facts: `standing.py:56` strips `miner=`; `gpu=` never parsed | `lib/standing.py:56`, `box-standing.sh:64` | operator-visible | 1 | fleet agent | MH/s, accepted and rejected, watts per box on the page (the HiveOS card) |
| Q91 | A transient ssh drop counts as death: `Box.run` returns 124 on timeout (`lib/box.py:49`), `check_one` marks dead on any rc (`standing.py:54`), the loop never asks the provider (`fleet.py:48-60` `refresh_ssh` unused); Vast proxy drops are known (`canary.sh:41`) | `lib/box.py:49`, `lib/standing.py:54` | operator-visible | 1 | fleet agent | provider status consulted before a dead check counts |
@ -355,7 +355,7 @@ The row for the ledger: **Q1 plus Q2, Q5, Q6, Q7, Q10**: every public and operat
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q96 | No log line on a finality flip: an operator tailing the log never sees the pause begin or end | `consensus/src/processes/finality.rs` (no pause line) | operator-visible | 1.5 | consensus-engineer | fast-time simnet prints `Finality: paused (<reason>)` and `Finality: resumed at index N`, one line per flip |
| Q97 | `finality_reason` says `paused` and nothing else: no share that left, no frozen-table hold, no expiry, no conflict reason (O-3.17) (the node half of Q1) | `finality.rs:1794-1807` | the project lead-visible (through every surface) | 1.5 (counted in Q1) | consensus-engineer | tonight's pause reads "paused: 42.7 percent of weight left the window; held by the table frozen at lock 6842 until DAA 216,402" |
| Q97 | `finality_reason` says `paused` and nothing else: no share that left, no frozen-table hold, no expiry, no conflict reason (O-3.17) (the node half of Q1) | `finality.rs:1794-1807` | the founder-visible (through every surface) | 1.5 (counted in Q1) | consensus-engineer | tonight's pause reads "paused: 42.7 percent of weight left the window; held by the table frozen at lock 6842 until DAA 216,402" |
| Q98 | No `/metrics`: finality_active, latest lock, exec tip, blocked, paid segments, peers, DAA as Prometheus gauges (the axum listener exists) | `igneum/exec/src/rpc.rs` | operator-visible | 3 | execution-engineer | `curl :26790/metrics` scraped by a Grafana dashboard committed under `infra/` |
| Q99 | Log level is start-only; no runtime change (monerod `set_log`, reth env filter) | `kaspad/src/args.rs:262-271` | operator-visible | 2 | consensus-engineer | an RPC or SIGHUP changes the level without a restart |
| Q100 | Clock-skew field in `getBlockDagInfo` and the local-ahead mirror case still owed after X19 | `docs/plans/node-changes.md:18-19` | operator-visible | 2 (after the `ledger-fixes-0311` merge) | consensus-engineer | `faketime -60s` gives one WARN and the field |
@ -395,8 +395,8 @@ So the trust today is one key, one builder (the box and the Mac, both the projec
| Id | Row | File or screen | Severity | Hours | Owner | Gate |
|---|---|---|---|---|---|---|
| Q3 | Signing and notarisation; first-run prompts (section 1) | `build-dmg.sh:118-122`, `windows.yml` | the project lead-visible | 6 plus certificates | app owner, the project lead | section 1 |
| Q80 | The Windows installer pipeline is dead: every GitHub-hosted job fails at start on billing; the installer and payload zip have no other builder, so 0.3.15 and every hotfix are Mac and HiveOS only until the project lead fixes Billing & plans, or the installer step moves to igneum-build-1 (Inno Setup under Wine, or a self-hosted Windows runner on PC 1) | `release-0.3.15.md` 6a, `.github/workflows/windows.yml`, `packaging/windows/build-installer.ps1` | the project lead-visible | 0 (the project lead: billing) or 4 (installer step on the box or PC 1 as a signed job) | app owner; the project lead | a green `windows.yml` run, or `tools/ship-app.mjs --dry-run` showing the fetch step satisfied from the new builder |
| Q3 | Signing and notarisation; first-run prompts (section 1) | `build-dmg.sh:118-122`, `windows.yml` | the founder-visible | 6 plus certificates | app owner, the founder | section 1 |
| Q80 | The Windows installer pipeline is dead: every GitHub-hosted job fails at start on billing; the installer and payload zip have no other builder, so 0.3.15 and every hotfix are Mac and HiveOS only until the founder fixes Billing & plans, or the installer step moves to igneum-build-1 (Inno Setup under Wine, or a self-hosted Windows runner on PC 1) | `release-0.3.15.md` 6a, `.github/workflows/windows.yml`, `packaging/windows/build-installer.ps1` | the founder-visible | 0 (the founder: billing) or 4 (installer step on the box or PC 1 as a signed job) | app owner; the founder | a green `windows.yml` run, or `tools/ship-app.mjs --dry-run` showing the fetch step satisfied from the new builder |
| Q81 | The ship tool pushes master whatever `--branch` says (the 19:30Z push of the release tree to master) | `tools/ship-app.mjs` commit step | operator-visible | 1 | app owner | `--branch X` pushes X; a test on a scratch repo |
| Q82 | The download page shows no hash for Windows and Mac, no OTA key fingerprint, no verify command; the stable aliases are not explained (Q50 covers the page; this row is the artefact side) | `miner.html:549-558`, `packaging/ota/publish-public.sh` | user-visible | 1 | site owner | `igneum-downloads.json` carries the fingerprint; the page prints it with `shasum -a 256` and `certutil -hashfile` lines |
| Q83 | The update key is one software key on one Mac, not in hardware, against G7's "held in hardware and its policy published" | `packaging/ota/README.md` "Keys", `docs/fud-ledger.md` G7 | operator-visible (security) | 3 (a YubiKey or Secure Enclave signer behind `igneum-ota-sign`, the policy page) | cryptographer, app owner | a release signed with the hardware key installs; the key policy is a public page; the software seed is retired |
@ -429,7 +429,7 @@ Forbidden strings (`site/forbidden-strings.txt`, 25 patterns) per shipped file,
| `site/wallet.html` | `dl\.igneum` | 1 | the wallet alias (`:371`) |
| `site/downloads.json` | `dl\.igneum` | 1 | the `base` field |
| `relay/ui.html` | `Hetzner` | 1 | the "Hetzner network" card title (`:377`); private console, rule does not apply, but one word |
| `app/igneum-app/ui/app.js` | `the project lead` | 1 | a code comment (`:270`); the ui bundle is outside every check (Q22) |
| `app/igneum-app/ui/app.js` | `the founder` | 1 | a code comment (`:270`); the ui bundle is outside every check (Q22) |
| `tools/community/discord-hooks.mjs` | `MacBook`, `\+0100`, `tailscale`, `ts\.net`, `dl\.igneum` | 1, 1, 1, 1, 2 | the guard's own regexes (`:51, 59, 61`), the read-only fleet URL (`:31`), and `ts\.net` matching "stats.net" in a template (`:255`): false positives, none posted |
| `pool/README.md` | `Hetzner`, `/opt/igneum`, `CLAUDE\.md` | 5, 2, 1 | the deploy section; not a site page, but it would fail the scrub if copied |
@ -448,7 +448,7 @@ Every other pattern is 0 on every file. Two-beat antithesis and aphorisms in shi
| Price | none anywhere; "pounds a day" once (`index.html:510`); Ember has a `priceNow()` with a devnet zero | | | | | | |
| Finality | "#N" or "paused" (`index.html:879`) | "#N, 3 min ago" plus the bar (`live.html:404, 512`) | none | "#N · 1 h ago" (`app.js:1350`) | "active" or "paused" | "checkpoint 6,842 at 55.x% of weight, 1 h ago, 93 voters" | none |
Row **Q85** (cross-cutting, 2 hours, site owner with the app owner): one shared formatter module (`site/lib/format.mjs`) for hash rate (1 decimal, unit ladder kH to PH), counts (`en-GB` separators, never `compact()` for the same quantity shown in full elsewhere), blocks per second (2 decimals, never inverted), difficulty (one spelling), IGN (4 decimals on tiles, 6 in tables), and one word for a vote identity ("vote key" on every surface); the app, hub, Discord and pool copy the same table as a Rust and a JS constant, with a test that renders the fixture values identically. Gate: the fixture renders byte-identical on all seven surfaces. Severity: the project lead-visible (he reads three of these side by side every evening).
Row **Q85** (cross-cutting, 2 hours, site owner with the app owner): one shared formatter module (`site/lib/format.mjs`) for hash rate (1 decimal, unit ladder kH to PH), counts (`en-GB` separators, never `compact()` for the same quantity shown in full elsewhere), blocks per second (2 decimals, never inverted), difficulty (one spelling), IGN (4 decimals on tiles, 6 in tables), and one word for a vote identity ("vote key" on every surface); the app, hub, Discord and pool copy the same table as a Rust and a JS constant, with a test that renders the fixture values identically. Gate: the fixture renders byte-identical on all seven surfaces. Severity: the founder-visible (he reads three of these side by side every evening).
### 4.4 Error states, one table
@ -484,7 +484,7 @@ Section 3.10 carries both in full. The two facts to carry forward: a fresh machi
- Lighthouse numbers are estimates from source; no browser run was allowed. The gate for every site row is a real run.
- The comparator claims are from memory; where a repository or page is named it is cited, otherwise approximate. The Stratum v2 reference was not cloned.
- The fleet and node sections depend on the gpu-fleet worktree and the vendor forks, read tonight; the fleet's own page was not opened.
- Hours are agent hours (the project lead's rule); Q3 and Q80 also need the project lead (certificates, GitHub billing).
- Hours are agent hours (the founder's rule); Q3 and Q80 also need the founder (certificates, GitHub billing).
## 7. Summary for the coordinator

View file

@ -111,13 +111,13 @@ The brief's `-pl` rows: the 5 October sweep on this card (`docs/plans/counter-as
| 65% | 400 (the floor) | 310.6 | 114.46 | 0.369 | 3,051 |
| 50% | 400 (clamped) | 302.4 | 109.20 | 0.361 | 3,050 |
`power.min_limit` is 400 W (read again by this job: min 400, max 600), so `-pl 200` and `-pl 250` cannot be set on a 5090; those rows are re-used, not re-run, and the ladder above adds what the cap does once the program carries work. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the project lead); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W.
`power.min_limit` is 400 W (read again by this job: min 400, max 600), so `-pl 200` and `-pl 250` cannot be set on a 5090; those rows are re-used, not re-run, and the ladder above adds what the cap does once the program carries work. Ember Tune's own 5090 steps did not run (its second engine never mined; re-run pending the founder); its idle readbacks on PC 1: 90.6 W idle, 2,505 MHz core, 14,001 MHz memory, limit 450 of 575 W.
The item 1 denominator moves: at the hash alone the 5090 is 290 W in the app (p95 290, max 295 W at 124 MH/s, the 5 October record on branch miner-eff: 2.34 microjoules) and 342 to 357 W in this bench (132 MH/s: 2.65 microjoules), not 326 W at 136.1 (2.40). The `f = 1` rows of chip-model-v3.md section 5.4 therefore read 5.0x (app) to 5.7x (bench) on GDDR7 against 5.1x, 7.3x to 8.3x on one HBM3 stack against 7.5x, 8.9x to 10.1x on eight against 9.2x: a move of 2 to 11 percent, inside the model's own margin.
## 6. The chip side, re-evaluated (approximate)
The `f = 1` chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at `k` times a GPU's marginal energy per op. The unit is now measured: 11 pJ per counted op, the 5090's marginal at its shipping clock (10.2 to 13.2 pJ on the three rungs under the cap, section 5); the model's 5.5 pJ (section 5.7, from TGP minus 326 W over the whole budget) is half that and is kept as the `k = 0.5` column, which is also about the M5 Max's 6.9 pJ. Four columns: `k = 1` (a chip core as good as the 5090's ALU on a random program: RandomX's argument and the project lead's test), `k = 1.5` (1.5x worse), `k = 0.5` (as good as the M5 Max's ALU, or the model's old unit), and `k = 0.3` (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 8x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the columns below compute `k` as written. Energy per hash = memory + N x 11 pJ x k; N in counted ops.
The `f = 1` chip of chip-model-v3.md section 5.4, at the memory's activate ceiling: GDDR7 (16 devices) 166.4 MH/s at 0.466 microjoules per hash, one HBM3 stack 83.6 MH/s at 0.321, eight stacks 0.262; those are memory, static and controller only. With the shadow the chip adds a core that runs the per-epoch random block at `k` times a GPU's marginal energy per op. The unit is now measured: 11 pJ per counted op, the 5090's marginal at its shipping clock (10.2 to 13.2 pJ on the three rungs under the cap, section 5); the model's 5.5 pJ (section 5.7, from TGP minus 326 W over the whole budget) is half that and is kept as the `k = 0.5` column, which is also about the M5 Max's 6.9 pJ. Four columns: `k = 1` (a chip core as good as the 5090's ALU on a random program: RandomX's argument and the founder's test), `k = 1.5` (1.5x worse), `k = 0.5` (as good as the M5 Max's ALU, or the model's old unit), and `k = 0.3` (a wide-SIMD fixed-datapath array at N5: the int32 datapath energy of section 5.1, 0.06 pJ per add and 0.52 per multiply at the op mix's 29 percent multiplies is 0.19 pJ, with a 2x pipeline overhead and about 8x for the register file, operand wires and the shared instruction fetch over a wide SIMD row, approximate; the floor a chip maker would claim). Section 5.7's "k = 1.5" column was computed as the chip being 1.5x BETTER (0.83 at N = 100,000 is 0.466 + 0.55 / 1.5); the columns below compute `k` as written. Energy per hash = memory + N x 11 pJ x k; N in counted ops.
ALU silicon for the core, at the 5090's 45.2 T op/s (what N = 330,700 at 136.1 MH/s needs): 22,600 32-bit lanes at 2.0 GHz. At N5 an int32 multiply-add lane with its register-file slice is about 0.002 mm^2 (approximate: 3,000 to 6,000 gates at about 0.3 square microns per NAND2-equivalent plus registers), so about 45 mm^2 of datapath, 60 to 100 mm^2 with SIMD control and operand networks, $25 to $40 of silicon at the sram-mirror.md yield model ($20,000 wafer, about $0.36 per mm^2 on a 128 mm^2 die; section 5 of that file); at 28 nm the same lanes are about 8x the area, 360 mm^2 of datapath and 500 mm^2 or more with control, a reticle-class die that a $5M to $30M controller project does not carry. Watts for the core at N = 330,700 and 136 MH/s: at `k = 1` 11 pJ x 45.2 T = 497 W (more than the whole 5090 draws for the same work, which is what `k = 1` means), at `k = 0.5` 249 W, at `k = 0.3` 149 W. At N = 100,000 the array is a third of that: 14,000 lanes, about 30 mm^2 at N5, 150 W at `k = 1`, 45 W at `k = 0.3`.

View file

@ -111,5 +111,5 @@ prototype values of spec 1.16, fixed at gate 1.
the 1,170 directly.
4. Monero's seven years do not price this device (ledger C13).
Decision at gate 1 (owner: the project lead): the cache size rule "exceeds what one die can hold, and grows", and the mixer
Decision at gate 1 (owner: the founder): the cache size rule "exceeds what one die can hold, and grows", and the mixer
cost multiplier, against the CPU verify gate.

View file

@ -373,7 +373,7 @@ Per tier at migration: a home miner on any card runs the updater and signs one `
| The testnet with no value | nothing; no asset | nothing | nothing | nothing | nothing |
| The external proving market paid in dollars | a job is a service; settlement in IGN on-chain with no operator custody is not a CASP activity; an operator that converts dollars to IGN is an exchange | the same; dealing or arranging if an operator sits between | money transmission if an operator holds funds | a dollar-settled service is a financial service only if the entity holds or exchanges | the same |
### 8.3 What Igneum writes into its genesis and public text now (not legal advice; counsel is engaged)
### 8.3 What Igneum writes into its genesis and public text now (not legal advice)
1. No issuer, no offeror, no sale: every coin is created as a block reward (MiCA Article 4(3)(b) language, verbatim in the litepaper).
2. The 2025/422 sustainability indicators (kWh a year from hash rate and measured uJ, intensity per transaction, mix by region when known) published by the project each era, so an EU CASP can list without asking.

View file

@ -2,7 +2,7 @@
7 October 2026, morning UK. Lane 4 of the final research programme ("the last mission"), worktree `igneum-wt-mission`. The question: given everything Igneum already has (the horizon one page, the frontier lane's sixteen ideas, the new-PoW lane's three schemes, class v5, finality in the proof, work-stake, the tail, vote-or-burn, the ladder, the pool design, the light client, the phone app), what is the NEXT genuinely new thing a GPU chain could do that none has. Every candidate here was checked against prior art by name and year, run through its own game theory with a model this lane ran (`invent_model.py`, section 0), reviewed in the Monero or Kaspa developer's voice, priced in agent hours with a first gate, and given its per-tier consequence and four persona verdicts. Nothing here touches the devnet. No em dashes. Figures from memory say approximate; figures from a source name it.
the project lead's rulings are applied as a filter before any idea is scored: miners are the security always; no stake and no coin bond; no holder penalised; no device bounty; no fixed-height activations; no calendar dates; no reliance on another chain; no dev fund; no privacy; nothing that becomes a token sale; the lottery and the proving stay separate; emission carries security with no end date. An idea that needs one of these is marked dead in one line.
The founder's rulings are applied as a filter before any idea is scored: miners are the security always; no stake and no coin bond; no holder penalised; no device bounty; no fixed-height activations; no calendar dates; no reliance on another chain; no dev fund; no privacy; nothing that becomes a token sale; the lottery and the proving stay separate; emission carries security with no end date. An idea that needs one of these is marked dead in one line.
## 0. Progress, method and the four voices

View file

@ -1,10 +1,10 @@
# The Igneum Mission
Written 7 October 2026, morning UK, by the mission coordinator (branch `mission`, worktree `igneum-wt-mission`) on the project lead's order of the same morning, verbatim: "i dont want to get stuck in a loop here where we keep finding and adding features and security so im going to give you one last mission, deep deep dive into the past and look far into the future, what has been done, what hasnt been done, what has been started but not finished, what can we invent and what can we reinvent to honestly make the perfect gpu network that will be noticed as the revolution that everyone is waiting for."
Written 7 October 2026, morning UK, by the mission coordinator (branch `mission`, worktree `igneum-wt-mission`) on the founder's order of the same morning, verbatim: "i dont want to get stuck in a loop here where we keep finding and adding features and security so im going to give you one last mission, deep deep dive into the past and look far into the future, what has been done, what hasnt been done, what has been started but not finished, what can we invent and what can we reinvent to honestly make the perfect gpu network that will be noticed as the revolution that everyone is waiting for."
This is the last research round. Its output is the closed list in section 2. After it the project stops researching and ships. Five lanes ran in parallel and each has its own file with sources and a verdict table: `past.md` (lane 1), `unfinished.md` (lane 2), `future.md` (lane 3), `invent.md` (lane 4), `reinvent.md` (lane 5). Every number below names its lane; the lane file names the source or the model. Hours are agent hours. Nothing on the live devnet was touched and no build or hash measurement was run for this document.
## 1. One page for the project lead
## 1. One page for the founder
**What has been done.** Of 31 GPU-mined chains since 2011, 8 lost the GPU lane to a chip, 7 closed it by choice, 10 died on price or a rental attack; the 6 still GPU-only pay USD 0.22 to 0.71 a day per RTX 3080 (lane 1). In 20 miner posts the wants were hardware that keeps its value, income without a cliff, a fair supply. Igneum answers the supply in full (no premine, no fund, no fee to any team), the hardware with a measured 2.1x chip bound and the N ladder, the cliff with a monthly glide and a 1 percent tail. Finality by 30 days of blocks, proving as the reward, the leave item and the signing bonus are built or approved for the testnet genesis.
@ -18,7 +18,7 @@ This is the last research round. Its output is the closed list in section 2. Aft
**The honest answer.** Igneum's revolution is the combination already built plus two new things: the finality object in the pocket, and the ladder shown everywhere. No shipped chain has the combination: a random-program GPU hash whose dataset is the chain, miner-only finality with no stake that 51 percent never reaches, that finality inside the execution proof, a work bond with no coin, a tail that holds a 24-hour attack at twelve days of emission in every year. Twelve items, about 290 agent hours, no new consensus rule beyond the three cut branches.
### Decisions owed from the project lead
### Decisions owed from the founder
The cryptanalysis entity and prize; the signing certificates (lane 5); the signed proving customer as a mainnet gate (lane 1); the H100 and A100 hash measurement on rented pods (lane 3).
@ -43,7 +43,7 @@ Twelve items, in build order. "In flight" names the lane or branch that holds it
### 2.1 The testnet genesis re-cut as approved (in flight, ship lane)
What it is: one cut with 18 decimals, `EmissionSchedule::TESTNET_1` (100 IGN a block, a monthly glide with a two-year half-life, a 90-day ramp from 10 percent, a 1 percent tail from year 11.4), and the switches on from genesis: proof verification in consensus, the leave item, the signing bonus at 1,000 bps, finality v3, the latency ladder at rung 0. the project lead approved it on 7 October 2026, 09:3x UK (ledger-decisions.md).
What it is: one cut with 18 decimals, `EmissionSchedule::TESTNET_1` (100 IGN a block, a monthly glide with a two-year half-life, a 90-day ramp from 10 percent, a 1 percent tail from year 11.4), and the switches on from genesis: proof verification in consensus, the leave item, the signing bonus at 1,000 bps, finality v3, the latency ladder at rung 0. The founder approved it on 7 October 2026, 09:3x UK (ledger-decisions.md).
Why it is right: lane 1's verdict rows 4 and 5 (the cliff and the premine) are answered by it in full; lane 3's tail row holds to year 200.
Evidence: tail-emission.md one page; vote-or-burn.md section 5; ledger-decisions.md 7 October.
Cost: 0 new hours; the ship lane's own plan.
@ -127,7 +127,7 @@ Order: ninth; before any phone trusts item 2's object.
What it is: lines added to `docs/plans/testnet-go.md` and the litepaper, each with its check. Gates: a signed proving customer before mainnet, with the weight of the Devnet 2 gate; per-tier income published at three prices (USD 0.005, 0.02, 0.10) before the testnet, per the consequences rule; the launch-week hash-origin report (key counts, pool shares, fleet shares) from the observer, daily for the first 90 days; the first 100 keys, the first 1,000 independent keys (the X5 definition), the first pool not run by the project, the first outside-reproduced benchmark, the first block from a card the project does not own; signed and notarised installers (Developer ID, Authenticode) with the measured first-share time per release. Text: the three wants on the front page and finality on page two; "a block pays its miner whether or not anyone buys a proof that day"; the eight regulatory sentences of lane 3 section 8.3 (no issuer, no sale; the sustainability indicators per era; no price language; the pool holds no balance; no privileged key), labelled "not legal advice; counsel is engaged".
Why it is right: lane 1's shape (income halves 60 to 120 days after a peak, under USD 0.10 per kWh within 6 to 18 months, 50 to 100 percent of hash gone in 90 days) and its rows 2, 4, 7 and 9; lane 5's install findings (SmartScreen warns on any file without reputation); lane 3's regulation rows (MiCA 4(3)(b) exempts block-reward assets; the UK regime from 25 October 2027 covers platforms and custody, not mining).
Evidence: past.md sections 3 and 4; future.md section 8; reinvent.md 3.1 and 4.1.
Cost: 15, plus the project lead's certificates and the customer's signature.
Cost: 15, plus the founder's certificates and the customer's signature.
Gate: every gate line in testnet-go.md with its check; the forbidden-strings check passes on the new text; the hash-origin report posts on Devnet 2 for seven days before the testnet go.
Order: tenth; text and gates can land any hour, and the customer gate is the one that takes longest.

View file

@ -2,7 +2,7 @@
Date: 7 October 2026, UK time. Lane 5 of the last mission. Written in the mission worktree; this file is the only output. Nothing was built, measured, posted or started; the live devnet was not touched.
the project lead's words that this lane holds: "revolutionise the mining itself"; "dopamine for GPU miners everywhere, a structure that makes the mining commodity fall in love with the project"; "the perfect gpu network that will be noticed as the revolution that everyone is waiting for". Bounds that this lane obeys: no premine, no dev fund, no treasury, no token sale, no airdrop, no referral paid in coins; every coin is mined under the rules that exist (the 80/20 split, the signing bonus of 8 of the 80 producer points, the proving pool); no stake; no device bounty; the three signalling thresholds (60, 90, 95); miners are the security always; gates, not dates; hours are agent hours.
The founder's words that this lane holds: "revolutionise the mining itself"; "dopamine for GPU miners everywhere, a structure that makes the mining commodity fall in love with the project"; "the perfect gpu network that will be noticed as the revolution that everyone is waiting for". Bounds that this lane obeys: no premine, no dev fund, no treasury, no token sale, no airdrop, no referral paid in coins; every coin is mined under the rules that exist (the 80/20 split, the signing bonus of 8 of the 80 producer points, the proving pool); no stake; no device bounty; the three signalling thresholds (60, 90, 95); miners are the security always; gates, not dates; hours are agent hours.
## 0. Progress
@ -148,8 +148,8 @@ What the comparators do: SmartScreen warns on any file that is not "well known a
| Step | Today | Proposal | Hours | Gate |
|---|---|---|---|---|
| Mac signature | ad hoc, quarantine strip | Developer ID plus notarisation and stapling in `build-dmg.sh`; the quarantine strip removed | 3 plus the project lead's USD 99 | a fresh macOS 26 machine opens the DMG's app with no Privacy and Security visit |
| Windows signature | none | Authenticode through Trusted Signing or an OV certificate, signed on the box as a job step; the sha256 and the signer's name on the download page (Q50, Q82) | 3 plus the project lead's purchase | a fresh Windows 11 machine runs the installer with no SmartScreen interstitial by the 1,000-download mark (reputation is use; before that the page shows the exact interstitial text and the two clicks) |
| Mac signature | ad hoc, quarantine strip | Developer ID plus notarisation and stapling in `build-dmg.sh`; the quarantine strip removed | 3 plus the founder's USD 99 | a fresh macOS 26 machine opens the DMG's app with no Privacy and Security visit |
| Windows signature | none | Authenticode through Trusted Signing or an OV certificate, signed on the box as a job step; the sha256 and the signer's name on the download page (Q50, Q82) | 3 plus the founder's purchase | a fresh Windows 11 machine runs the installer with no SmartScreen interstitial by the 1,000-download mark (reputation is use; before that the page shows the exact interstitial text and the two clicks) |
| Firewall prompt | UAC 20 to 50 s in, unexplained | the Cards screen names it and the step runs after `setup_done` (Q14) | 1 | the prompt never appears before a screen that says it will |
| Antivirus | nothing | a "What your antivirus may say" line on `/miner#get` with the engine names and the exact action, and the submission to Microsoft's SmartScreen review on every release (the Learn page's submission route) | 1 | the line is on the page; every release's binaries are submitted the day they ship |
| Time to the first share | not measured end to end on a fresh machine | measured on a fresh Windows 11 VM and a fresh Mac for every release by the Devnet 2 gate: download to first accepted share or block, with the clock, the sync and the dataset build as three timed steps | 2 | "first share within 10 minutes of download in 9 of 10 fresh Windows installs", published on `/evidence` |
@ -364,7 +364,7 @@ Written in copy law (short sentences, no antithesis, no aphorism).
| Ember, Earnings | "£0.00 earned", mined IGN never shown (Q18) | IGN a day, blocks and IGN in 30 days, shards and IGN, a typed price, share of network, weight rank, days, signing streak | 6.5 | every line names its RPC field on hover; a view test per line |
| Ember, Prove | state line "0 assigned · 0 proven · 0 paid" | the shard card "your card proved shard N of block M, 1.15 IGN" | 1 | a card on the first paid record on Devnet 2 |
| Ember, Settings | pool field designed, not built (`pool.md` section 7) | the pool field, the terms box from `welcome`, the dev-fee line reading "off in pool mode: the pool's 1 percent is the same fee" | 3 | a member connects from the app and shows the terms |
| Ember, first run | two prompts, one unexplained (Q3, Q14) | signed and notarised; the firewall step named on the Cards screen | 7 plus the project lead's certificates | 3.1 gates |
| Ember, first run | two prompts, one unexplained (Q3, Q14) | signed and notarised; the firewall step named on the Cards screen | 7 plus the founder's certificates | 3.1 gates |
| Wallet | one account, no finality state on the address page (Q33, Q37) | the first-transfer card with the five state words; the address page with the latest lock | 3 (plus Q9's 6) | a devnet transfer walks the five words on screen |
| Site `/miner` | "Install. Start. The card mines and proves." plus feature lines | the first-hour timeline (download, the one prompt, first share, first block, first payout) with the measured times from the Devnet 2 gate | 2 | the times on the page equal the gate's log |
| Site `/live` | miners with a block in 10 minutes | the weight leaderboard with days mined and signing presence; the 10-minute view as a toggle | 3 | top 100 by weight; a stopped key falls one place a day |

View file

@ -1,6 +1,6 @@
# The last mission, lane 2: started and not finished, across the industry
7 October 2026. Lane 2 of the final research programme. What the industry started and did not finish, group by group, with one verdict per row: Igneum already has the finished version, Igneum could finish it in hours, or Igneum should leave it. Every figure from a source carries its URL and access date in section 9; every figure from memory is labelled approximate. Hours are agent hours (the project lead's rule: Claude-side work takes hours, never weeks). Nothing here is a token sale.
7 October 2026. Lane 2 of the final research programme. What the industry started and did not finish, group by group, with one verdict per row: Igneum already has the finished version, Igneum could finish it in hours, or Igneum should leave it. Every figure from a source carries its URL and access date in section 9; every figure from memory is labelled approximate. Hours are agent hours (the founder's rule: Claude-side work takes hours, never weeks). Nothing here is a token sale.
What this file does not repeat: the ASIC chip history (`docs/analysis/asic-resistance-history.md`), the useful-work verdict and the stored-state candidate (`new-pow.md` sections 3, 6, 7), the ranked frontier ideas (`frontier.md` section 0, 3.10, 3.11), and Igneum's own pool design (`docs/plans/pool.md`, `docs/spec/09-pool-protocol.md`). Those are cited, not restated.
@ -257,7 +257,7 @@ Ember's state is from `polish.md` 3.1 and 3.10: a Rust engine runs `igneumd` and
| Pool mode | Every miner | Design only | Behind; the share sidechain of 5.3 is the finish |
| Finality on the surface | Nobody | Nothing during a pause (Q2) | Behind, 3 hours, and nobody else has the concept |
The honest line: Ember is the first app to ship the thing the table says nobody shipped (node, wallet, GPU miner, own key, signed updates, no custody), and it loses to a 2018 Honeyminer on the sixty seconds between download and first share. The first-run fixes are one Developer ID and one Authenticode certificate (Q3, a the project lead decision) plus the Q80 installer builder; the rest of the behind rows are 11 hours.
The honest line: Ember is the first app to ship the thing the table says nobody shipped (node, wallet, GPU miner, own key, signed updates, no custody), and it loses to a 2018 Honeyminer on the sixty seconds between download and first share. The first-run fixes are one Developer ID and one Authenticode certificate (Q3, a the founder decision) plus the Q80 installer builder; the rest of the behind rows are 11 hours.
## 7. Verdict table
@ -280,7 +280,7 @@ The honest line: Ember is the first app to ship the thing the table says nobody
| E | Encrypted member transport | SV2 Noise | No (plain TCP, Q67) | 8 | Finish |
| E | A pool-as-contract with sampled share verification (SmartPool) | Nobody since 2017 | No | 0 | Skip (1.35 ms per share makes it worse than it was for Ethereum) |
| F | Node, wallet and GPU miner in one app, own key, signed updates | Nobody; Ember | Yes | 0 | Already done |
| F | Sixty seconds from download to first share | Honeyminer did it in 2018 | No (three prompts, no Windows build) | 6 plus certificates (Q3); 0 or 4 (Q80) | Finish; the certificates are the project lead's |
| F | Sixty seconds from download to first share | Honeyminer did it in 2018 | No (three prompts, no Windows build) | 6 plus certificates (Q3); 0 or 4 (Q80) | Finish; the certificates are the founder's |
| F | Earnings in IGN, per-card faults, an hour of history, a finality notice | HiveOS, lolMiner | No (Q18, Q13, Q16, Q2) | 11 | Finish |
| F | Payout in something a gamer can spend (Salad's gift cards) | Salad | No, and the chain says "nothing is bought or sold on devnet" | 0 | Skip until there is a market |

View file

@ -1,6 +1,6 @@
# The prover floor: why a 12 GB card cannot prove on SP1 6.8.1's GPU server, and the patch
5 October 2026, 22:00 UTC on (the project lead: "execute if it will solve the issue"). Branch `prover-floor`
5 October 2026, 22:00 UTC on (the founder: "execute if it will solve the issue"). Branch `prover-floor`
(worktree `igneum-wt-prover-floor`). The measured facts this starts from: `docs/plans/proving-v1.md` and the
bench-log entry "proving v1" (branch proving-v1): the GPU server holds 13.9 GB for an empty shard, 20.4 GB at the
adopted v1 shard, 28.3 GB flat from 20 M to 60 M cycles, and no environment knob moved the floor. Every figure

View file

@ -1,6 +1,6 @@
# Prover tiers on real cards: the memory matrix of the patched SP1 GPU server, measured on rented GPUs
6 October 2026, from 11:50 UTC (the project lead: "rent all you need, absolute overkill", "get as many GPUs as you need to properly
6 October 2026, from 11:50 UTC (the founder: "rent all you need, absolute overkill", "get as many GPUs as you need to properly
test everything swiftly"). Branch `gpu-fleet`, tools in `tools/fleet/`, raw logs per instance under
`~/Desktop/fleet/<instance>/` (the sampler csv, every point's host log and results JSON, the miner log). This file
replaces the 5090-allocation rows of `docs/analysis/prover-floor.md` with the cards themselves. A row that is not here

View file

@ -1,11 +1,11 @@
# Proving methods: why the prover needs 14 GB, what else exists, and how a 12 GB card gets to prove
5 October 2026, from the project lead at 22:05 UTC: "if this doesn't enable 12 GB cards, then do a full deep research task on proving
5 October 2026, from the founder at 22:05 UTC: "if this doesn't enable 12 GB cards, then do a full deep research task on proving
and see if there are different methods." "This" is the prover-floor agent's patch of SP1's GPU server (branch
`prover-floor`), running tonight. This document is research and reading, not measurement: every number of ours is from
`docs/bench-log.md` with its entry named; every claim about another system cites its repository file, its documentation
page or its paper, or is labelled approximate. Status words follow `docs/spec/00-overview.md` 0.2. Day estimates follow
the project lead's rule of 3 October 2026: hours of agent time, never weeks.
The founder's rule of 3 October 2026: hours of agent time, never weeks.
The facts this starts from (bench-log, "proving v1", 5 October 2026; `docs/plans/proving-v1.md`; `docs/analysis/amd-proving.md`):
@ -351,7 +351,7 @@ the tier it moves, in one sentence, the day it is taken.
| Consequence | Action | Owner |
|---|---|---|
| The gate card has never run a prover here | get a 4070 or 3060 into the measurement loop this week; until then every 12 GB figure stays approximate | coordinator; the project lead for the card |
| The gate card has never run a prover here | get a 4070 or 3060 into the measurement loop this week; until then every 12 GB figure stays approximate | coordinator; the founder for the card |
| Route A's gate | the prover-floor agent's rows (asked for by message tonight); if under 11 GB, `provedefault.rs` gains the 12 GB prove-only and 16 GB mine-and-prove tiers and the rig installer one server per card | prover-floor agent, then the proving engineer |
| Route D's measurement | RISC Zero 3.0.6 at po2 19 and 20 on PC 2 (CUDA) and on this Mac (Metal), the same shard statement run natively: memory, time per segment, lift and join, receipt size | prover-floor agent (PC 2); a Mac measure job for Metal |
| The two-family node items (record version selects the verifier, fresh chain at a version change, one family per block) | spec 7.8 gains the three rules when route D starts; nothing changes before | execution engineer |

View file

@ -402,7 +402,7 @@ prove nothing. The static check is the guard, as it is for the dataset mask.
4. The lever that moves the named chip is the M16 mixer multiplier: x2 to 1.2x, x4 to 0.6x against the measured
5090 rate, at 0.8 to 4.8 ms of verification per warp against the 10 ms gate. Decision 2 should price that
against the gate on the 2019-class core (O-1.14) rather than layer 3.
5. If the project lead keeps layer 3 for a reason outside this analysis: ship the host contract in the pack, add the four
5. If the founder keeps layer 3 for a reason outside this analysis: ship the host contract in the pack, add the four
vector items of section 6 (consecutive units with a forced collision, the wrap launch, the warp-count-independent
fingerprint, the contract text), and run the CUDA and OpenCL twins of section 7.2 on the PCs before the class
becomes a genesis rule.

View file

@ -229,7 +229,7 @@ Option E is the only one that makes the mirror a multi-die part and it costs eve
genesis (at the headline density; the lower-bound density would ask for 4 GiB and 3.2 s), which fails the spirit of
the 10 ms verify gate (the fill is once a day, but a light node joining pays it on every day it syncs across).
Decision for the project lead, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a
Decision for the founder, at gate 1: A, B, C, D or E above, together with M16's mixer multiplier. Nothing here changes a
vector today: the cache size is a prototype value of spec 1.16 and the growth rule would be a new sentence in 1.13.3.
## 8. Why the latency bound is the property to lean on (citations behind the plan's rule)

View file

@ -250,7 +250,7 @@ Attack, before and after (ignored test `measure_m15_attack_before_and_after`, re
- Before (the pre-fix order, replayed by calling the engine with the header's own unvalidated day seed): 50 cold 256 MiB builds, 10,595 ms, and the honest day's entry is evicted (`KEEP` was 3).
- After (the new order): 0 builds, all 50 rejected (`TimeTooOld` or `UnexpectedHeaderDaaScore`) in 14 ms total. The live day stays resident.
Tests: kaspa-pow `--features igneum-pow` 8 pass (engine smoke, `one_build_per_seed_pair_under_contention`, `live_days_survive_off_day_builds`, `build_queue_is_bounded`, index and live-day helpers, shared-engine, stub); kaspa-consensus header_processor `cheap_checks_run_before_the_pow_engine` pass; kaspa-p2p-flows `pow_guard` 2 pass; the full kaspa-consensus release suite otherwise unchanged.
M16 Metal note (R3.5, cheap reconfirmation only): the Mac `--inline-dataset` shortcut at the 256 MiB cache, 256 MiB dataset, under the same heavy load, ran honest 91.7 Mhash/s against inline 5.29 Mhash/s (inline about 17x slower); this is noisier and slower than the idle-machine figures already in `proto-metal/MEMHARD.md` (10x slower at a 256 MiB dataset, 4.8x at 1 GiB), because the inline kernel is compute-bound and the machine was loaded. The 64 MiB on-die-SRAM emulation M16 wants (inline kernel with a 64 MiB cache, `cacheLog2Words = 24` in `proto-metal/main.swift`) is the RTX 5090 run reserved for the project lead's PC, as R3.5 states; it is not done here and the Mac number above does not price a die.
M16 Metal note (R3.5, cheap reconfirmation only): the Mac `--inline-dataset` shortcut at the 256 MiB cache, 256 MiB dataset, under the same heavy load, ran honest 91.7 Mhash/s against inline 5.29 Mhash/s (inline about 17x slower); this is noisier and slower than the idle-machine figures already in `proto-metal/MEMHARD.md` (10x slower at a 256 MiB dataset, 4.8x at 1 GiB), because the inline kernel is compute-bound and the machine was loaded. The 64 MiB on-die-SRAM emulation M16 wants (inline kernel with a 64 MiB cache, `cacheLog2Words = 24` in `proto-metal/main.swift`) is the RTX 5090 run reserved for the founder's PC, as R3.5 states; it is not done here and the Mac number above does not price a die.
Not done: the real-engine daemon RPC run (honest blocks need GPU-mined pow, so the measurement used the equivalent validate path with `skip_proof_of_work`); the 64 MiB-cache inline kernel on the 5090; the chain-derived day seed by DAA score (spec 01 section 1.12, still the timestamp-day devnet rule); the VDF epoch seed and finality.
## 3 October 2026, difficulty controller: devnet record, simulator, Igneum dual-lane rule, 3-node CPU test network (consensus-engineer)
@ -418,7 +418,7 @@ Not changed: the pgas table magnitudes (prototype), `B_p` = 30 M (prototype). Op
Machine for the reproduction: Apple M5 Max, 64 GiB, Darwin 25.6.0, load average 2 to 147 (other agents' builds and, during R1, another agent's Metal worker on the same GPU); everything at `nice -n 19`. Binaries: HEAD `proto-metal/main.swift` built with `swiftc -O` into the scratchpad (465,529 bytes, the same size as `proto-metal/igneum-bench`), `vendor/igneum-node-diff/target/release/igneumd` and `igneum-miner` (22:38 and 22:17 BST, the `difficulty` worktree pair; the miner's `Seeder` and worker protocol are the same code as HEAD and as the Windows build 745d41ef). Private networks on 127.0.0.1 ports 27500 to 27562, appdirs under `/tmp/igneum-decay-test`, all stopped afterwards. Full write-up: `docs/analysis/hashrate-decay-2026-10-03.md`; proposed fix: `docs/analysis/hashrate-decay-2026-10-03.patch` (not applied; `git apply --check` passes against `vendor/igneum-node`).
PC data (`node tools/logs.mjs <run_id> --all`, STATUS lines deduplicated by timestamp, per-interval rates from consecutive cumulative figures): segment 22:57 to 23:04 UTC, nvidia-1: 40 jobs in the first 30 s then exactly 32 per 30 s for 12 intervals at 17.5 to 18.4 MH/s wall while the printed cumulative figure fell 22.18 to 18.13; nvidia-8 (started 4.7 s later) printed a rising 16.80 to 17.71. Segment 22:23 to 22:57 UTC (epoch 2, DAA 8,474 to 10,513): per-identity gap between jobs 0.098 s to 0.330 s per 0.68 to 0.81 s job, inside-jobs rate rising 28.7 to 34.7 MH/s, wall falling 24.6 to 20.6 MH/s, card total 197 to about 165 MH/s; at the 22:57 epoch boundary the gap returned to 2% and the difficulty held (84.5M to 83.0M).
Code audit: nothing allocated per job survives the job in `proto-cuda/host.cu`, `proto-opencl/host.c` or `proto-metal/main.swift` serve loops (tables in the analysis); the miner's only per-job growth is time in `Seeder::seeds_for` (memo keyed by `(epoch, sink)`, one `getBlock` RPC per block from the sink to the epoch start on every miss, 1,274 to 3,313 calls on the PC). `cudaDeviceSynchronize` at the default schedule spins one thread per worker (the project lead's 6.2% per process); the hot-swap working tree sets `cudaDeviceScheduleBlockingSync` and swaps `clFinish` for `clWaitForEvents`.
Code audit: nothing allocated per job survives the job in `proto-cuda/host.cu`, `proto-opencl/host.c` or `proto-metal/main.swift` serve loops (tables in the analysis); the miner's only per-job growth is time in `Seeder::seeds_for` (memo keyed by `(epoch, sink)`, one `getBlock` RPC per block from the sink to the epoch start on every miss, 1,274 to 3,313 calls on the PC). `cudaDeviceSynchronize` at the default schedule spins one thread per worker (the founder's 6.2% per process); the hot-swap working tree sets `cudaDeviceScheduleBlockingSync` and swaps `clFinish` for `clWaitForEvents`.
Metal runs (STATUS every 30 s; "gap" = 1 minus wall over inside, per interval): R1 control, epoch 0, genesis bits 0x1d100000, 308 s (cut by the 22:21:37 UTC SIGTERM of every process of this session): first interval 25.18 MH/s alone on the GPU, then 14.0 to 14.4 MH/s in every interval after another agent's worker joined at 25 s, gap 0 to 2%, worker RSS 56.8 MiB flat. R2 walk reproduction, 900 s: `skip_proof_of_work` node pumped to DAA 4,000 (one-second timestamps, difficulty held at 76.8M), one identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s: inside 27.0 to 27.7 MH/s in all 29 intervals; wall 22.3 to 24.3 (gap 12 to 18%, walk 400 to 700), 25.9 to 27.3 (gap 0 to 4%), 18.0 to 21.0 (gap 25 to 35%, walk 700 to 1,000); miner CPU 0 to 1% in the quiet phase, 11 to 21% in the last. R3 one worker at difficulty 2^25 (Kaspa sampled rule, genesis bits held), 600 s: 30.51 wall / 30.72 inside, 1,091 jobs, 224 blocks, 53 to 56 jobs per 30 s throughout. R4 eight workers at 2^25: 29.38 / 29.45 summed (3.32 to 4.38 each), 1,053 jobs, 242 blocks, 7 jobs per 100 s per identity in every interval, worker CPU 0.0 to 0.6%, RSS 46 to 57 MiB. R5 one worker at 2^31: 37.01 / 37.75, 1,324 jobs, 4 blocks, flat. R6 eight workers at 2^31: 36.67 / 36.75 summed (4.29 to 5.55 each), 1,314 jobs, 5 blocks, flat. (R5 and R6 ran a different epoch-0 program from R3 and R4, 112 loads per hash, hence 37 against 30.5 MH/s.)
Side findings: the `difficulty` worktree's node panics at `consensus/src/processes/difficulty.rs:431` ("Work should not exceed 2**192") when fed 85 blocks/s with wall-clock timestamps under the Igneum dual rule (a pump artefact, logged for the consensus-engineer); `skip_proof_of_work` nodes still log "PoW rejected ... by igneum-lottery-v1-bound" for every block they accept. Not done: the fix applied and measured on the PC (the acceptance figure is a flat gap at DAA 10,800 with eight identities); the OpenCL event wait checked on the AMD driver; a unit test of `seeds_for` (the client is concrete).
@ -526,7 +526,7 @@ Binaries: `target-integration/release/{igneumd 40,463,680 B, igneum-miner 7,916,
## 4 October 2026, generator version 2 adopted: exact load count, fresh-source loads, program acceptance; every vector re-cut, three workers re-checked, 20,000-program census, devnet-v4 binaries rebuilt (cryptographer)
Machine: Apple M5 Max, idle at the start (load 3), rustc 1.99.0 (rustup), Swift 5.8.1, builds at nice 10. Decision (the project lead, this morning): adopt the census's generator rule before any public vector ships. Rule as implemented (`igneum-pow/src/generator.rs`, `src/accept.rs`, spec 01 sections 1.4.2, 1.4.3 and 1.4.6; mirrored in `proto-metal/main.swift` as `generateProgramV2` and `acceptProgram` because the Metal worker derives its program from the seed itself): G1 exactly 16 load slots, a uniform subset of instructions 1..63 drawn first by partial Fisher-Yates, the other 48 ops from the ten non-load weights (sum 75); G2 a load's source is drawn from the registers other than `dst` written by an earlier instruction and not read by a load since; R (a) no cyclically stale load source, (b) every register has an injecting write, (c) 64 units at base nonces from `SplitMix64(FNV-1a-64("igneum-accept/" || seed words LE))` on the seed-keyed closed-form dataset at 2^28 words with init words = seed words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes; a rejected candidate is replaced by `seed_words_from_bytes(seed || k_le32)`, `k = 1, 2, ...`, 32 consecutive rejections a consensus fault. Every pack carries `generator 2`, the attempt and the program id `FNV-1a-64("igneum-program/" || 2_le32 || seed words LE || attempt_le32)`; `igneum-pow` is 0.2.0 and the node's engine reports `igneum-lottery-v2-bound`. The crate is now the pack source (`igneum-pow export`); the Swift exporter is the Metal cross-check. Version 1 stays as `generate_v1` for the census and MEMHARD.md levers; its vectors are retired.
Machine: Apple M5 Max, idle at the start (load 3), rustc 1.99.0 (rustup), Swift 5.8.1, builds at nice 10. Decision (the founder, this morning): adopt the census's generator rule before any public vector ships. Rule as implemented (`igneum-pow/src/generator.rs`, `src/accept.rs`, spec 01 sections 1.4.2, 1.4.3 and 1.4.6; mirrored in `proto-metal/main.swift` as `generateProgramV2` and `acceptProgram` because the Metal worker derives its program from the seed itself): G1 exactly 16 load slots, a uniform subset of instructions 1..63 drawn first by partial Fisher-Yates, the other 48 ops from the ten non-load weights (sum 75); G2 a load's source is drawn from the registers other than `dst` written by an earlier instruction and not read by a load since; R (a) no cyclically stale load source, (b) every register has an injecting write, (c) 64 units at base nonces from `SplitMix64(FNV-1a-64("igneum-accept/" || seed words LE))` on the seed-keyed closed-form dataset at 2^28 words with init words = seed words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes; a rejected candidate is replaced by `seed_words_from_bytes(seed || k_le32)`, `k = 1, 2, ...`, 32 consecutive rejections a consensus fault. Every pack carries `generator 2`, the attempt and the program id `FNV-1a-64("igneum-program/" || 2_le32 || seed words LE || attempt_le32)`; `igneum-pow` is 0.2.0 and the node's engine reports `igneum-lottery-v2-bound`. The crate is now the pack source (`igneum-pow export`); the Swift exporter is the Metal cross-check. Version 1 stays as `generate_v1` for the census and MEMHARD.md levers; its vectors are retired.
Packs regenerated (`proto-cuda/packs/`): `igneum-genesis` and `igneum-hourly` (closed form), `igneum-genesis-mh` (memory-hard, day 2026-10-03, cache FNV unchanged `48c4f5bf24166b2e`), and new `igneum-devnet-v4-epoch0` (epoch seed = devnet genesis hash `edc4fa84...fb07`, day bytes `igneum-day/20730`, cache FNV `448274a57f508cbc`). `igneum-genesis` attempt 0, program id `bcc1248b10cc90f2`, op mix `load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1`, lane 0 at base 0 `42246ba99fc58e4f`, lane 31 `b08446b1f2de7793`; devnet pack id `4be132dd1f2ff270`, lane 0 `285a83011e7ac3fc`. Bound vectors re-cut (`igneum-pow/README.md`: H zero, nonce 0 gives `746c567b090acf6a`). `cargo test --release`: 39 of 39 (28 unit, 11 pack).
@ -545,7 +545,7 @@ Not done: no NVIDIA or AMD hardware has run a version 2 pack (the RTX 5090's 192
## 2026-10-04 finality floor 2/3: the total-weight floor raised from 17/30 to two thirds, simulator A to L re-run, attack scenarios 6A and 6B on a three-node, six-voter network (cryptographer)
Decision of 4 October 2026 (the project lead, O-3.15): a lock needs two thirds of all 30-day weight, and finality pauses whenever less than two thirds of that weight is connected and signing; the chain continues on proof of work meanwhile and the node reports it. Spec 3.3, 3.3.1, 3.7, 3.9, 3.11 rewritten; litepaper Finality and "What Igneum does not claim" updated; ledger F2, F9, F16, F18 restated and F21 added (the window bound of attack scenario 6A).
Decision of 4 October 2026 (the founder, O-3.15): a lock needs two thirds of all 30-day weight, and finality pauses whenever less than two thirds of that weight is connected and signing; the chain continues on proof of work meanwhile and the node reports it. Spec 3.3, 3.3.1, 3.7, 3.9, 3.11 rewritten; litepaper Finality and "What Igneum does not claim" updated; ledger F2, F9, F16, F18 restated and F21 added (the window bound of attack scenario 6A).
Node. Branch `devnet-v4` of `vendor/igneum-node` (worktree `vendor/igneum-node-v4`, from `dc749905`), commit `6457ca95`, two files: `consensus/core/src/finality.rs` (`FLOOR_NUM / FLOOR_DEN` 2/3, was 17/30; the Q3 arithmetic as `FinalityParams::{quorum_met, floor_met, locks}`, both comparisons inclusive) and `consensus/src/processes/finality.rs` (`lock_test` calls it). Build `CARGO_TARGET_DIR=target-integration nice -n 10 cargo build --release -j 6 -p kaspad --features kaspad/igneum-pow`, 3 min 17 s on a machine at load 3 to 13 (another agent's igneum-pow rebuild and the live devnet running). Tests `cargo test --release -j 6 -p kaspa-consensus-core -p kaspa-consensus -- finality`: 7 of 7 in consensus-core including the new `floor_is_two_thirds_of_total_and_inclusive` (4 of 6 locks, 3 of 6 does not, 2 of 3 locks, 67 of 100 locks, 66 does not, 57 does not; the total test implies the active test at every participation; a 3/3 side never locks whatever the other side's participation decays to), 2 of 2 in consensus (`no_certificate_while_the_window_is_filling` unchanged). The live devnet (26610, 26611, 26640, 26641, 28640, the seed relay on 26680 and `observer.mjs`) was never touched.
@ -577,7 +577,7 @@ Notes. (1) The v4 node logged "PoW rejected ... by igneum-lottery-v1-bound" abou
## 4 October 2026, proving v0 on the RTX 5090: first GPU proof of an Igneum block (WSL2, SP1 6.8.1 cuda)
Machine: the project lead's Windows 11 PC, RTX 5090 (32,607 MiB, driver 617.14), 16 cores and 45 GB visible to WSL2 Ubuntu 24.04, mining
Machine: the founder's Windows 11 PC, RTX 5090 (32,607 MiB, driver 617.14), 16 cores and 45 GB visible to WSL2 Ubuntu 24.04, mining
paused. Package `proving/windows-wsl2` (SETUP-PROVER then PROVE-BLOCK), host `igneum-prove-host` built with the `cuda` feature,
`SP1_PROVER=cuda`, sp1-gpu-server 6.8.1 on device 0. Fixture `block-78-increment` (chain 4463, 2 transactions, 10 accounts).
Run id `prove-<pc>-20261004-084838`, log intake id 10154.
@ -773,7 +773,7 @@ button and the held miner all showed; with the real clock the HTTPS source read
## 4 October 2026, first machine on the Igneum Miner app: PC 2's RTX 5090 at 118 MH/s, via Setup.exe
the project lead's second PC (a clone of the first; the app's per-install machine id `1ccfe586` keeps its keys apart), installed from the
The founder's second PC (a clone of the first; the app's per-install machine id `1ccfe586` keeps its keys apart), installed from the
runner-built `Igneum-Miner-Setup-0.3.0.exe` (unsigned, SmartScreen "run anyway"), the one-click package: prebuilt NVRTC
worker, no toolchain on the machine. First attempt sat at "waiting for peer": the PC's clock was 62 s slow after a power cut
and `igneumd` rejected every relayed block ("the block timestamp is too far into the future"; the 10-s skew bound from the
@ -794,7 +794,7 @@ Candidates on the live replay (join window, 3 seeds, std of log difficulty / lan
Cost (synthetic set seed 7, v1 / v2): x50 settled 61.7 / 65.5 s (standard under 90), /50 628 / 753 s (worst gap 32 / 62 s), epoch +-30% settled 144 / 143 s, hop10 211 / 200 s, polluted overshoot 0.180 / 0.089, warm-ups equal, steady std 0.038 / 0.049 with blocks-per-minute CV 0.135 / 0.130. Attacks (seeds 7 to 9, v1 / v2): greedy hopper at most +1.5% / +2.5%, with a 60 s dwell -4.0% / -1.5% at 100%; pulsed rental -96.4% / -96.3% with weight per hash 0.262 / 0.262; forger drift +0.4 to +1.1% / -0.8 to +0.5% (worst seed 2.7% / 1.5%); short-lane oscillation gain 3.75 / 3.29; epoch games 0.0 to +0.7% / 0.0 to +0.4%; polluted window settled 287 to 329 s / 288 to 331 s; base profiles 3-seed up50 154 / 150 s, down50 762 / 822 s, epoch30 88 / 88 s, hop10 245 / 233 s, polluted 75 / 76 s, steady std 0.042 / 0.053.
Implementation (devnet-v4): `difficulty_v2_activation_daa` in `Params` (every network `u64::MAX`), `OverrideParams`, `override_params`, the daemon's file parser (prints the height), `SampledDifficultyManager` (new field, `reference_window(daa_score, epoch_blocks, activation)`), `REF_WINDOW_V2 = 600`, `IgneumInputs.k_ref`; `infra/fast-time/override-60x.json` carries the field as never. `cargo test --release -p kaspa-consensus --lib difficulty`: 15 pass (12 of 3 and 4 October plus `reference_window_switches_at_the_activation_height`, `v2_reference_window_follows_a_step_inside_the_epoch_where_v1_eases_into_it`, `v1_and_v2_agree_in_a_steady_epoch`); `-p kaspa-consensus-core --lib params`: 7 pass (`override_params_carry_the_difficulty_v2_activation`, the fast-time file test extended). Build 2 min incremental for `igneumd` and `igneum-miner`.
Test network (`sim/difficulty/testnet_v2.py`, 3 nodes on 29600 to 29622, the 60x file with the devnet epoch, genesis bits 2^16, activation 900 on nodes 1 and 2, node 3 without it; CPU miners A from 0, B from minute 4, off at 19, back at 23): node 1 reached DAA 900 at 1,022 s; node 3 rejected the first v2 block ("difficulty of 520437997 is not the expected value of 520406991"), banned its peer and stayed at DAA 900 (901 headers, a prefix of node 1's 1,472); nodes 1 and 2 agreed on every header and the sink. Under v2 the leave eased 6,589 to 5,972 over 180 s (std 0.036, no peak), the rejoin hardened 6,154 to 8,312 within 60 s and held within 3%. The v1 phase is not readable: the load swung the CPU miners' delivered hash rate 2x on its own (difficulty fell 40% after B joined). Record `records/testnet-v2-2026-10-04.csv`. Repeat on a quiet machine, 30 minutes.
Rollout: only `igneumd` changes (the Mac build, `infra/cross/build-linux.sh` for the seed and the Hetzner nodes, the Windows package for PC 1's node); every node of a chain needs the same `"difficulty_v2_activation_daa": N` in its override file before the height or it forks off there. First the 12 Hetzner nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, PC 1's node, in that order, by the project lead. A new network sets 0.
Rollout: only `igneumd` changes (the Mac build, `infra/cross/build-linux.sh` for the seed and the Hetzner nodes, the Windows package for PC 1's node); every node of a chain needs the same `"difficulty_v2_activation_daa": N` in its override file before the height or it forks off there. First the 12 Hetzner nodes on their own chain (N = current DAA + 1,800, restart one by one, a miner joins inside an epoch, no flips after the height), then the devnet with N about two hours ahead: observer node, seed, Mac node 1, PC 1's node, in that order, by the founder. A new network sets 0.
Not done: a quiet-machine test-network run; the DAG model's red blocks and the Mac's log series; v2 with +-500 ms stamp jitter.
## 4 October 2026, the observer stored nothing for 78 minutes, then 7,022 blocks in two minutes
@ -856,7 +856,7 @@ first mismatch and are not a rate.
## 4 October 2026, shard proving on the RTX 5090: a full shard compressed in 10.9 s, a two-shard block aggregated in 2.2 s, all verified
Machine: the project lead's PC 2 (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the
Machine: the founder's PC 2 (RTX 5090, 32,607 MiB; WSL2 Ubuntu 24.04, SP1 v6.8.1 with the cuda feature, the toolchain of the
4 October morning run in `~/igneum-prove`), the Igneum Miner app (machine id 1ccfe586) stopping its miners for the
run. Delivered as the signed job `shard-benchmark` (`app/igneum-app/src/jobrun.rs`, `packaging/ota/publish-jobs.sh`),
which runs `prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4"` and reports every RESULT line to the
@ -1067,7 +1067,7 @@ first four minutes; the identity's votes on checkpoints 1202 and 1203 were accep
discrete card takes part in finality. Machine id 37ba0461 in the console; app log run `win-37ba0461-20261004-185342`.
## 4 October 2026 (evening), finality rule v3: the frozen weight table (F21) and the certificate fold (F22), simulator, unit tests, fast-time 3-node network with 300-ms links (finality engineer)
the project lead, 4 October 2026 evening: "we need to fix these serious issues before making things public". Both fixes sit behind one height switch, `finality_v3_activation_daa` (default never on every network, set by the override file like `difficulty_v2_activation_daa`), on branch `finality-fixes` of the node (worktree `vendor/igneum-node-finality`, from the proving head `8c0cff15`). Spec 03 Q4 (fold) and Q5 (frozen table), 3.3.1, 3.7 items 2 and 9, 3.10, 3.11; ledger F21 and F22 "Fix built, pending rollout"; the devnet plan in `docs/plans/finality-v3-rollout-devnet.md`. The live devnet was never touched; every network below ran on ports 29700 to 29799, suffix 970.
The founder, 4 October 2026 evening: "we need to fix these serious issues before making things public". Both fixes sit behind one height switch, `finality_v3_activation_daa` (default never on every network, set by the override file like `difficulty_v2_activation_daa`), on branch `finality-fixes` of the node (worktree `vendor/igneum-node-finality`, from the proving head `8c0cff15`). Spec 03 Q4 (fold) and Q5 (frozen table), 3.3.1, 3.7 items 2 and 9, 3.10, 3.11; ledger F21 and F22 "Fix built, pending rollout"; the devnet plan in `docs/plans/finality-v3-rollout-devnet.md`. The live devnet was never touched; every network below ran on ports 29700 to 29799, suffix 970.
**F22, what was wrong.** The cloud logs of the healthy stretch 11:45 to 14:00 UTC (212 indices, 12 miners; `tools/finality-attacks/vote-timing.py`, output in `infra/cloud-devnet/results/2026-10-04/f22-vote-timing.md`): the node builds a certificate the instant the votes it holds meet Q3, median 1.24 s (p99 1.71 s) after the first node determined the checkpoint, with 7 to 10 of 12 signers (mean 8.27); 10.24 votes had been issued by then on average (two in flight: the miner's 1-s poll, the 250-ms gossip pump per hop, up to 289 ms RTT) and the last of the 12 was issued median 1.45 s, p90 2.36 s after the first determination. A 1-s hold after the first build would have carried all 12 votes at 192 of 212 indices; the other 20 are miners 01, 06 and 11 down together for 10 minutes (indices 377 to 396, the hop.sh restarts), an outage, not lag. There is no cut-off to lengthen: the fix is a second round. Presence needs nothing, since the block reading of Q2 credits a late vote once any block carries it.
@ -1094,7 +1094,7 @@ the project lead, 4 October 2026 evening: "we need to fix these serious issues b
| split50, v3: the same cut | neither side locks during the split; heal; locking resumes on one chain; 0 conflicting certificates | 0 / 0 / 0 new locks during the 150 s (the v2 control locked at 126 s, so the frozen table held side B for the checkpoints of the last 24 s; the frozen table would have expired at 240 s); n0 redialled 72 s after the gate reopened; all three nodes resumed at index 7 and reached 13 inside the heal window; 0 conflicting certificates; 0 disagreeing locked indices; every post-heal LOCKED line names the frozen lock and its fraction (81 to 94% of the frozen table) | PASS |
| split70, v3: 4/2 keys, the 4 side at 70% of weight (shares 0.175 x 4 against 0.15 x 2), 150 s | the 4 side locks during the split, the 2 side does not; 0 conflicts | 4 side: 4 new locks, the first 30 s after the cut; 2 side: 0; heal: all three at 17; 0 conflicting certificates; 0 disagreeing indices | PASS (6B at 70/30; exactly 4/6 is a knife edge under both rules, simulator row above) |
**What remains uncertain.** (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node Hetzner network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. the project lead asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from `daa(C_i) + depth`, a few seconds late at worst. (6) Binaries, all from `finality-fixes` 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native `fe982a1d...` (verified running), Linux `7c100fc2...` (cargo-zigbuild, 34 min, not run on a Linux host), Windows `cc1d1001...` (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with `push-inputs.sh --no-deploy` into a scratch folder and NOT deployed (plan 7a).
**What remains uncertain.** (1) The v3 split's hold was observed over the last 24 s of a 150-s split (the control crossed at 126 s); a longer split under the 240-s expiry (say 200 s) would show more held checkpoints, and the "held by the frozen table" line is logged at debug, which the runs did not enable. (2) No cloud rehearsal: the 12-node Hetzner network was destroyed at 15:30 UTC, so the 95% target is shown on three nodes with emulated 300-ms links and on the cloud logs' arithmetic, not on the cloud topology itself; the rollout plan names the re-creation and the partition experiment to run first. (3) The frozen reference is the node's own highest lock on the chain, not the certificate carried in C_i's past, so two honest nodes can test one checkpoint against tables 30 s apart; in a connected network those tables differ by a minute of blocks, under a partition both are pre-split, and no run showed a disagreement, but it is a property argued, not proved. (4) The price: a sudden departure of a third or more now pauses finality for a full window (30 days on mainnet) instead of 1.4 to 10 days; a gradual one costs nothing. The founder asked for the pause over the fork; the number is stated in spec 3.7 item 2. (5) The fold clock is in memory: a restarted node folds from `daa(C_i) + depth`, a few seconds late at worst. (6) Binaries, all from `finality-fixes` 6aa69a45, hashes and checks in the rollout plan's section 2: Mac native `fe982a1d...` (verified running), Linux `7c100fc2...` (cargo-zigbuild, 34 min, not run on a Linux host), Windows `cc1d1001...` (mingw, 12 min 28 s, the v2 exe's DLL set, cannot run here); the Windows payload inputs were staged with `push-inputs.sh --no-deploy` into a scratch folder and NOT deployed (plan 7a).
## 4 October 2026, miner performance: variant racing (Metal worker on the M5 Max; the RTX 5090 job is ready, not run)
@ -1553,7 +1553,7 @@ Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converg
## 5 October 2026 (evening), the 9070 XT on the eGPU: why 17.9 MH/s, and what moved
PC 1 (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string `3683.0 (PAL,LC)`, platform `OpenCL 2.1 AMD-APP (3683.0)`). Branch `opencl-rdna4`. the project lead: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).
PC 1 (ae432dc7, Windows 11, Ryzen 7 9800X3D with its gfx1036, RTX 5090 on CUDA), an AMD Radeon RX 9070 XT (gfx1201, RDNA 4) in a Sonnet Breakaway Box 850T5 over USB4, Adrenalin 26.9.2 (OpenCL driver string `3683.0 (PAL,LC)`, platform `OpenCL 2.1 AMD-APP (3683.0)`). Branch `opencl-rdna4`. The founder: "the hashrate is low" (17.9 MH/s with one worker; two workers on the card earlier gave 8.9 and 9.4).
**Before, from PC 1's own app log** (`node tools/logs.mjs win-ae432dc7-20261005-181046`, the miner's STATUS line for the card `amd:1:gfx1201`, 2^21-nonce jobs): `hash=17.82 MH/s wall (17.83 MH/s inside jobs) ... idle=0.3%`. Wall equals inside, so the host loop (template fetch, job line, read-back, scan) costs nothing measurable; the dispatch itself is slow. The worker's `ready` line: `exchange 0` (local memory: AMD lists `cl_khr_subgroups` and no shuffle extension), `batch 4194304`, `dataset-log2 28` (1 GiB), device `[1] gfx1201` on the 3683.0 platform, `AMD wavefront width 32`. The same card was listed again as `[3] gfx1201` on the older platform `3652.0` (the 32.0.21042 driver's OpenCL registration is still present after the update): that is the two-worker run.
@ -1589,12 +1589,12 @@ Reading: at the dataset size the card delivers about 2.5 G random 4-byte reads p
| Card | 1024 MiB chase at 4,096 lanes | 1024 MiB chase ceiling | indep x8 ceiling | ceiling / 128 = hash ceiling | measured hash rate |
|---|---|---|---|---|---|
| RX 9070 XT, eGPU over USB4 | 2.64 G/s, 1,552 ns | 2.42 to 2.68 G/s | 2.42 G/s | 18.9 to 20.9 MH/s | 18.0 to 18.1 MH/s (bench), 17.8 (app) |
| RTX 5090, PCIe 5 x16, contended | 9.09 G/s, 451 ns | 16.4 to 18.0 G/s | 16.2 to 16.7 G/s | 128 to 141 MH/s | 127 MH/s (app, the project lead), 139.7 alone (M11) |
| RTX 5090, PCIe 5 x16, contended | 9.09 G/s, 451 ns | 16.4 to 18.0 G/s | 16.2 to 16.7 G/s | 128 to 141 MH/s | 127 MH/s (app, the founder), 139.7 alone (M11) |
| Apple M5 Max, Apple OpenCL | 2.10 G/s, 1,949 ns | 3.41 to 3.49 G/s | 3.45 to 3.47 G/s | 26.6 to 27.3 MH/s | 27.9 Mhash/s (README, Apple OpenCL) |
Reading: on all three cards the hash runs within a few percent of 1/128 of the card's dependent random-read ceiling, which is what a 128-load program should do; the probe is a good model of the hash. The 5090 does 6.6x the random reads of the 9070 XT for 2.8x the rated bandwidth (1,792 against 640 GB/s, vendor figures): the rest is access granularity and DRAM behaviour on random 4-byte reads, which the kernel cannot change.
**Power, heat, fans and clocks, measured** (branch `opencl-rdna4-telemetry`; the project lead watched the 9070 XT at 90% usage with its fans barely turning and the app had no AMD reading, the MH/W line came from nvidia-smi only; a new helper `proto-opencl/gpu-telemetry.c` reads ADLX on Windows and the amdgpu sysfs on Linux. Job `tele-measure-1`, 20:27:45 to 20:29:41 UTC, both cards mining in the app, nothing touched: `igneum-gpu-telemetry -l 5` (sha256 `703cf69c…a9c69b`) and `nvidia-smi --query-gpu=index,name,power.draw,temperature.gpu,fan.speed,clocks.mem,clocks.gr,utilization.gpu -l 5` side by side, the app's `hash_now` every 5 s; `node tools/jobs.mjs tele-measure-1`):
**Power, heat, fans and clocks, measured** (branch `opencl-rdna4-telemetry`; the founder watched the 9070 XT at 90% usage with its fans barely turning and the app had no AMD reading, the MH/W line came from nvidia-smi only; a new helper `proto-opencl/gpu-telemetry.c` reads ADLX on Windows and the amdgpu sysfs on Linux. Job `tele-measure-1`, 20:27:45 to 20:29:41 UTC, both cards mining in the app, nothing touched: `igneum-gpu-telemetry -l 5` (sha256 `703cf69c…a9c69b`) and `nvidia-smi --query-gpu=index,name,power.draw,temperature.gpu,fan.speed,clocks.mem,clocks.gr,utilization.gpu -l 5` side by side, the app's `hash_now` every 5 s; `node tools/jobs.mjs tele-measure-1`):
| Card | Samples | Watts (mean, min to max) | Temperature | Fan | Memory clock | Shader clock | Busy | Hash (mean of 24) | MH/W, measured |
|---|---|---|---|---|---|---|---|---|---|
@ -1632,7 +1632,7 @@ Reading: the kernel is the same 116.0 ms on both paths (18.08 MH/s pure kernel,
**A second defect found on the way: the pack export race.** PC 1's app log since its 19:02 UTC restart (`node tools/logs.mjs win-ae432dc7-20261005-190232`): `worker error: error 0 pack packs\devnet: the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT` at 19:07:03, 19:07:19 and 19:08:07, so the 9070 XT was not mining at all in the app while this entry was written (my job `rdna4-serve-1` at 18:43 hit the same folder in the same state). Cause, from `app/igneum-app/src/engine.rs` `prepare_worker`: one thread per card, each running `igneum-miner export-pack` into the one folder `packs\devnet`; across an epoch change the two exports interleave and the folder keeps one epoch's `program.h` with the other's `seeds.txt` until the next export. Fix on this branch: a process-wide mutex around both export sites (`EXPORT_LOCK`); the second export rewrites the same pack. Not measured in the app yet: it ships with the branch.
**Answer to the project lead.** The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.
**Answer to the founder.** The 9070 XT does 2.5 G random 4-byte reads per second from its memory for this access pattern, and the hash needs 128 of them, so about 19 MH/s is this card's ceiling for the current program class, on any slot; it was running at 92% of that. The eGPU link cost 6% per job through the read-back, now removed (17.87 against 16.88 MH/s inside jobs standalone). The duplicate platform that halved it to 8.9 + 9.4 is folded away. The pack race that stopped it is serialised. Nothing else in the worker's control moves the number: the next step for this card is the program class itself (fewer, wider loads per hash would favour AMD's 64-byte lines), which is a consensus question, not a worker one.
## 5 October 2026 (night), Ember Tune: the two-knob efficiency tune, the fleet prior, and what PC 1 could measure tonight (miner-community-lead)
@ -1658,7 +1658,7 @@ Branch `ember-tune` (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH
| RX 9070 XT (bus 98, present again) | `tune 1 ... gmax 0 gmax_range -500 1000 plimit 0 plimit_range -30 10 factory 1 ok` | the helper's clock range is an OFFSET from stock in MHz, not a ceiling: a probe reading it as a 1,000 MHz maximum would have asked for `--set-gmax 900`, an overclock. Fixed at 054e041: an offset range closes the clock knob (until the stock clock is known) and the power ladder runs on the percent scale bounded by the range, so the 9070 XT's plan is 100, 90, 80, 70% (the -30 floor), 4 steps |
| Radeon(TM) Graphics (integrated) | `tune 0 ... gmax - ... factory 0 ok` | no manual tuning: measure only, and it is off by default anyway |
**Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the project lead's word, engine 25113f52..., PC 1 on 0.3.11):** the project lead answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's `Set-Content -Encoding utf8`, which adds a UTF-8 BOM; the engine's JSON parser refuses it, `Settings::load` fell back to defaults (no payout address, no cards), the engine logged `[error] no payout address` and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (`gmax 0`, `plimit 0`). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (`job finished; the miners restart` at 07:56:50Z, both miners up by 07:57:04Z, `mining` at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (`RESULT TUNE scratch settings: address ..., cards N, first bytes ...`), and the CI check fails any playbook writing JSON with `Set-Content -Encoding utf8`. The re-run needs one more click on the prompt.
**Run 2, 6 October 2026, 07:21 to 07:56Z (job ember-tune-pc1-2, elevated on the founder's word, engine 25113f52..., PC 1 on 0.3.11):** the founder answered the one prompt; the installed app stopped its miners at 07:21:16Z; the second engine ran for the whole 35-minute budget at "waiting, 0.00 MH/s" and no step ran. Cause: the playbook wrote the engine's copy of settings.json with PowerShell 5.1's `Set-Content -Encoding utf8`, which adds a UTF-8 BOM; the engine's JSON parser refuses it, `Settings::load` fell back to defaults (no payout address, no cards), the engine logged `[error] no payout address` and never started a miner. Run 1's scratch log carried the same line the night before. Readbacks, idle both times: the 5090 at 90.6 W before and 69.9 W after (2,505 then 2,407 MHz core, 14,001 MHz memory, limit 450 W of 575), the 9070 XT at factory (`gmax 0`, `plimit 0`). Nothing set on either card. The installed app's runner released the miners-stopped hold by itself on the failed exit (`job finished; the miners restart` at 07:56:50Z, both miners up by 07:57:04Z, `mining` at 07:57:29Z): mining paused 36 min 13 s. Fix 8273494: the copy is written without a BOM, the address is read back and the job fails within seconds if it is empty (`RESULT TUNE scratch settings: address ..., cards N, first bytes ...`), and the CI check fails any playbook writing JSON with `Set-Content -Encoding utf8`. The re-run needs one more click on the prompt.
**Dry run 3, 6 October 2026, 14:56 to 15:02Z (job ember-dryrun-pc1-3, unelevated, no prompt, measure only; engine from ember-tune 07d5a72, kit sha256 36b522c9...):** the first measurement engine on PC 1 that mined. Both cards, one 60 s row each at the installed app's 80% cap, clocks unlocked, rate = the worker's STATUS wall rate, draw = nvidia-smi every 5 s:
@ -1669,7 +1669,7 @@ Branch `ember-tune` (54ff1bc), docs/plans/ember-tune.md. Every card tuned for MH
Nothing set; the installed app's miners back after 350 s. Why every earlier run (5 and 6 October, runs 1 to 4 and dry runs 1 and 2) read its copied settings as defaults, measured on PC 1 (collect ember-acl-2): the engine's own start locks its app folder with `icacls /inheritance:r /grant:r <user>:F`; cutting the folder's inheritance propagates down, the non-inheritable grant gives the children nothing, so a file COPIED in before the start (settings.json, machine-id, wallet.json) is left with no access entry and its owner cannot read it (`ReadAllText`: access denied), while the engine's own files written after the lock inherit fine, which hid it for a day. A first fix with `(OI)(CI)F /T` left the file empty too: `/T` re-applies `/inheritance:r` to each file after the propagation and an `(OI)(CI)` entry on a file is inherit-only. The right form is the inheritable grant without `/T` (07d5a72). Consequence for every tier on Windows: nothing changes for the installed app (its files were always its own); any tool that drops files into the app folder before the app starts (an installer's seed, a migration, a support script) was unreadable to the app until now and is readable from 0.3.13 on.
**Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the project lead's click, PC 1 on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct):** the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:
**Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the founder's click, PC 1 on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct):** the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:
| Cap | Limit | MH/s | W | MH/W | GPU C |
|---|---|---|---|---|---|
@ -1679,12 +1679,12 @@ Nothing set; the installed app's miners back after 350 s. Why every earlier run
| 70% | 403 W | 127.38 | 310.9 | 0.410 | 65 |
| 60% (floor 400 W) | 400 W | 127.38 | 311.3 | 0.409 | 65 |
Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an `-lgc` lock persists until `-rgc` or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a `--sweep` engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the `after` snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the project lead switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (`rgc`), the tune runs unelevated through the task.
Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an `-lgc` lock persists until `-rgc` or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a `--sweep` engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the `after` snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the founder switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (`rgc`), the tune runs unelevated through the task.
Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on PC 1 is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.
## 5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)
Branch `readwidth` (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in `docs/plans/read-width.md`. Nothing here changes consensus: every class sits behind `igneum-pow --class` and the default class is generator version 2 byte for byte (`igneum-pow/tests/packs.rs` passes on the four pinned packs after every commit). Question (the project lead, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).
Branch `readwidth` (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in `docs/plans/read-width.md`. Nothing here changes consensus: every class sits behind `igneum-pow --class` and the default class is generator version 2 byte for byte (`igneum-pow/tests/packs.rs` passes on the four pinned packs after every commit). Question (the founder, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).
**What a class does** (`igneum-pow/src/generator.rs` `LoadClass`, `verify::fold_words`, the three emitters): a load of W words reads the W-word-aligned address `(src AND MASK) AND NOT (W - 1)` and folds every word into `dst` (`x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]`); W = 1 is the lottery hash exactly (`w4` = pack `bcc1248b10cc90f2`). A mix class draws W per load with one extra `below(100)` roll per instruction. A scratch class `scr<k>k<kb>` turns `k` of the 16 memory slots into read-modify-writes of a 16-byte slot of the lane's share of a `kb` KiB per-warp scratch (kernels run persistent warps, one per block or work-group; a slot reads as a seed-and-base fill until the unit writes it, behind a per-unit tag). Program ids carry the class. Dependent chain and 32-lane unit unchanged.
@ -1862,7 +1862,7 @@ Commands: `IGNEUMD=vendor/igneum-node/target-txgossip/release/igneumd IGNEUM_MIN
## 5 October 2026 (night), the SP1 CPU prover on PC 1 beside the miners, and the backend survey: no zkVM proves on AMD (amd-prove agent)
the project lead, 22:50 BST: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: `docs/analysis/amd-proving.md`. The survey (SP1 v6.8.1 and `dev` 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA.
The founder, 22:50 BST: "test proving on the amd card?" and "can we test proving on mac?". The analysis with the backend table and the tier consequences: `docs/analysis/amd-proving.md`. The survey (SP1 v6.8.1 and `dev` 318dd530 of 28 Sep 2026, RISC Zero, Jolt, OpenVM, ICICLE, sppark; every claim cites a file or page there): on 5 October 2026 no zkVM proves on an AMD GPU; Apple silicon has RISC Zero's shipped Metal prover and ICICLE's Metal backend; SP1, the prover here, is CPU-only off NVIDIA.
Machine: PC 1 (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 cores and 46,994 MB visible to the VM, the Igneum Miner app 0.3.9 mining on the RTX 5090 and the RX 9070 XT throughout (the 5090 at 89% mean utilisation, 59 to 70% minimum, from a 1-s `nvidia-smi` sampler under the run: the job never touched a card). Signed `run` job `cpu-prove-pc1-small2` (`tools/amd-prove/pc1-cpu-prove.ps1`), 20:49:00Z to 20:59:49Z: the hosted package `igneum-prove-wsl2-pv1b.zip` (sha256 df50dee5...) built WITHOUT the `cuda` feature (6 s warm; the first job `cpu-prove-pc1-small` built it cold in 126 s), `--mode id` the pinned pair (shard `0x2b1a81cb...`, aggregator `0x474678f3...`, pinned 2026-10-05T16:20:38Z), `SP1_PROVER=cpu`, `--mode shard --shard 0` under `/usr/bin/time -v`. Log: `node tools/jobs.mjs cpu-prove-pc1-small2 --all`.
@ -1878,7 +1878,7 @@ Reading, and the consequences (CLAUDE.md, every number). Doubling the cycles add
## Counter ASIC 2.0, the numbers
5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on PC 1 and PC 2, RX 9070 XT on PC 1's eGPU), the decisions taken under the project lead's delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: `docs/plans/counter-asic-2-public.md`.
5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on PC 1 and PC 2, RX 9070 XT on PC 1's eGPU), the decisions taken under the founder's delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: `docs/plans/counter-asic-2-public.md`.
**Program class v3 (the devnet, activation by height switch `program_class_v3_activation_daa`)** = class v2's 128 x 4-byte loads, the era draw of the table layout and the working-set windows (layers 4 and 8), the cache growth rule (layer 6, option C: the cache doubles when the dataset doubles), the mixer at x8 (M16's multiplier), reserve family R1 (integer matrix, switched off) and the epoch length as a signalled reserve parameter (layer 9, 3,600 DAA s until a 90% signal). Not adopted on the measurements: wider reads (layer 1), the per-load width mix (layer 2), the per-warp write scratch (layer 3), the hot table (layer 5).
@ -1963,7 +1963,7 @@ Inputs, all RTX 5090 (PC 2), SP1 6.8.1 cuda: a full shard at the provisional `S_
Reading. The card count is the sum of card-seconds of work per block-second, rounded up, with no slack for the exclusive window, the relay or a card's idle gaps; the devnet's own numbers tonight (one card, 1.4 to 1.6 shards a minute, 2.4 to 4.7% of blocks) are the first row. Two levers, both measured tonight: the loop (a shard's carriage through export, cut and a 12-s key setup is 25 s on top of a 7-s proof; the host's `--mode aggregate` and `--mode chain` hold one key setup per process and the prover loop should do the same, the 0.3.11 item in the plan) and the card's other job (a mining card proves 3 to 4x slower than an idle one, `chain-pc2-pv1c` against 4 October; the prover's cost to mining is 4%). A fleet of 18 mining 5090s, or 6 proving-only ones, covers an empty-block chain at 1 block/s through the chain mode; the mandatory rule waits for the measured share to reach one, not for these rows.
### The 12 GB requirement (the project lead, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs
### The 12 GB requirement (the founder, 20:1xZ: "make sure we can prove on 12gb cards"): the GPU memory peak against SP1's knobs
Job `memsweep-pc2-pv1` (`tools/proving-v1/pc2-memory-sweep.ps1`), PC 2's RTX 5090 (32,607 MiB), the miners STOPPED by the job and the live prover switched off (its `sp1-gpu-server` would otherwise be the one the client connects to), every row: the server killed first, a 1-s `nvidia-smi memory.used` sampler, one `--mode compressed --shard 0` run of the pv1 host (`/opt/igneum-pv1`, SP1 6.8.1 cuda, `sp1-gpu-server` 6.8.1), 20:19 to 20:25Z. The knobs are the environment the GPU server inherits from the host process (`sp1-core-executor-6.8.1/src/opts.rs`: `SHARD_SIZE`, `ELEMENT_THRESHOLD`, `HEIGHT_THRESHOLD`, `MINIMAL_TRACE_CHUNK_THRESHOLD`, `TRACE_CHUNK_SLOTS`; `sp1-prover-6.8.1/src/worker/config.rs`: the `SP1_WORKER_NUM_*` and `*_BUFFER_SIZE` counts, defaults 4 core workers, 8 recursion prover workers). Idle card before the sweep: 1,732 MiB.
@ -2069,7 +2069,7 @@ Totals: 228 of 228 PASS where expected, 3 of 3 FAIL where built in. Reading: the
## 5 October 2026 (night), epoch length as an era parameter (Counter ASIC 2.0, layer 9): the Mac's compile-ahead per program
Branch `ca2-epoch`, worker "ca2-epoch"; design and the per-card table in `docs/plans/epoch-length.md`. Question (the project lead: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (`tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh`). `proto-metal/igneum-bench` built from this branch with `swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal` under a build slot.
Branch `ca2-epoch`, worker "ca2-epoch"; design and the per-card table in `docs/plans/epoch-length.md`. Question (the founder: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (`tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh`). `proto-metal/igneum-bench` built from this branch with `swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal` under a build slot.
Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/epoch9`; version 2 generator, 128 loads per hash), each generated and compiled at run time (`makeLibrary` from source plus `makeComputePipelineState`), dataset 2^28 words, one 2^20 batch and one verify warp per program:
@ -2158,7 +2158,7 @@ cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.
11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core
(2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The
build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below;
the 9070 XT is OWED (PC 1 is the project lead's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier
the 9070 XT is OWED (PC 1 is the founder's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier
already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to
94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the
same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache;
@ -2232,7 +2232,7 @@ Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5
Reading: the 5090 holds to 150,800 ops (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W cap that the control never reaches (342 to 357 W) and that binds from 102,100 ops up: the SM clock falls from 3,037 to 1,834 MHz and at 330,700 ops the card is compute-bound at the capped clock (28.6 T counted op/s, the 45.2 T budget scaled by the clock). The 5 percent point at 431 W is about 210,000 ops. The marginal ALU energy at the shipping clock, read on the three rungs under the cap: 10.2 to 13.2 pJ per counted op, twice the 5.5 pJ the chip model assumed. The 64-instruction block runs 3.5 percent above the control here too. Clock rows (`-lgc`): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them. Power-cap rows: the 5 October sweep's (floor 400 W, so `-pl 200` and `250` cannot be set; the cap never binds at the control).
**RX 9070 XT (PC 1, ae432dc7)**: OWED (PC 1 is the project lead's desk and not released today); the OpenCL kernels are in every pack.
**RX 9070 XT (PC 1, ae432dc7)**: OWED (PC 1 is the founder's desk and not released today); the OpenCL kernels are in every pack.
Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (`sh256x27`) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 holds its rate at its 431 W cap (350 W at the control: per watt 0.81x, measured), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the `f = 1` chip's edge per joule falls from 1.6x to 0.9x against the M5 Max and from 5.6x to 2.1x against the 5090 at `k = 1`, where `k` is the chip core's energy per op over the 5090's measured 11 pJ: the number that decides the item. Verdict: GO as a class v4 candidate at N = 100,000 (`mx8+sh256x27`), subject to the 9070 XT row and the gates; NO-GO above 130,000 or with a block over 256 instructions. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.
## 6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs
@ -2283,7 +2283,7 @@ Reading of the Mac rows. The run-to-run spread is under 4% on every row. The dot
Reading of the 5090 rows. Every row is bit-exact, `mm8` included, so the m8n8k16 fragment layout of the CPU reference (PTX ISA 9.4 section 9.7.16.5.3) is the layout the hardware uses. The `alu` chain reads 7,941 G steps/s here against 8,754 through OpenCL event time on 5 October: a CUDA event pair around a 0.54 ms kernel carries about 0.05 ms of launch, which also compresses every ratio toward 1 (approximate: the ratios are the card's at the 2% level, not better). On this card every candidate costs more than the reference chain, unlike Apple: the 5090 runs the two-register add-xor-rotate chain at one IMAD and one funnel shift per step, and the three-register candidate chains pay their glue. Against the live `rotr` (1.32), the candidates read: `andn` 0.95x, `shl` 0.96x, `shr` 0.97x, `perm` 0.98x, `sel` 1.00x, `popc` 1.14x, `shfla` 1.16x (the same as the live xor shuffle, 1.49: the indexed shuffle costs NVIDIA nothing extra), `bfe` 1.17x, `clz` 1.23x. `bfe.u32` and the C form cost the same to the nanosecond (0.835 ms), so the compiler emits the same code for both and no single-instruction bit-field extract is in play on this architecture (not checked by cuobjdump; the equal times are the evidence). `dp4a` reads 1.16x (1.17x on 5 October). `mm8` is the most expensive row on NVIDIA too (2.43x the reference: one tensor-core mma per warp per dependent step, latency-bound), which supports its place at the end of the reserve on the honest-card side as well as on the chip side.
**RX 9070 XT (PC 1, ae432dc7), OpenCL**: OWED. PC 1 is the project lead's desk and not released today (the brief's rule); the OpenCL twin of the probe (`__builtin_amdgcn_*` paths for `v_bfe_u32`, `v_perm_b32`, `v_bcnt_u32_b32`, `v_cndmask_b32`, `ds_bpermute_b32`) is the next job on that card.
**RX 9070 XT (PC 1, ae432dc7), OpenCL**: OWED. PC 1 is the founder's desk and not released today (the brief's rule); the OpenCL twin of the probe (`__builtin_amdgcn_*` paths for `v_bfe_u32`, `v_perm_b32`, `v_bcnt_u32_b32`, `v_cndmask_b32`, `ds_bpermute_b32`) is the next job on that card.
Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at `W_new` = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live):
@ -2348,7 +2348,7 @@ Cause: `ProvingState.paid_wei: u128` and serde_json `to_value` (1.0.151, `value/
## 5 October 2026 (night), aggregation cost on the RTX 5090: what a per-block aggregation spends and what each lever gives (proving engineer, agg-cost)
the project lead, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on PC 2's 5090 while the card mined (`chain-pc2-pv1c`, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch `agg-cost` (worktree `igneum-wt-agg-cost`, from `proving-v1` 219517f). Host changes (statement untouched, `elf/` untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, `--mode chain --save-shards` (every shard's compressed proof written next to the results, so `--mode aggregate` re-runs the same proofs under other settings). Jobs: `agg-cost-pc2-1` (21:01:20Z to 21:25:11Z, `tools/proving-v1/pc2-agg-cost.ps1`, the package `igneum-prove-wsl2-aggcost.zip` fetched by `fetch-prove-aggcost` 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to `/opt/igneum-aggcost`, the live `/opt/igneum` untouched, `--mode id` the pinned pair) and `agg-cost-pc2-2` (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through `/tmp/sp1-cuda-0.sock` and carry its own environment; `gpu_server_before running=0`) and ON again at the end. Fixtures: four consecutive live blocks cut from PC 2's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on PC 2 throughout.
The founder, 5 October 2026: "fix everything else in the numbers tonight". The number under test: the chained segment aggregation cost 9.6 to 9.7 s a block on PC 2's 5090 while the card mined (`chain-pc2-pv1c`, the entry above), 2.2 s on 4 October with the card to itself. Target: under 3 s a block, the miner's slowdown of the prover under 1.5x, the proof statement unchanged. Branch `agg-cost` (worktree `igneum-wt-agg-cost`, from `proving-v1` 219517f). Host changes (statement untouched, `elf/` untouched): the aggregation's stdin build timed apart from the prove call, the deferred-proof count and the SP1 knobs in the RESULT lines, `--mode chain --save-shards` (every shard's compressed proof written next to the results, so `--mode aggregate` re-runs the same proofs under other settings). Jobs: `agg-cost-pc2-1` (21:01:20Z to 21:25:11Z, `tools/proving-v1/pc2-agg-cost.ps1`, the package `igneum-prove-wsl2-aggcost.zip` fetched by `fetch-prove-aggcost` 20:55:39Z, built in WSL2 against the live target dir in 5 s, installed to `/opt/igneum-aggcost`, the live `/opt/igneum` untouched, `--mode id` the pinned pair) and `agg-cost-pc2-2` (21:34:00Z, the same script). The live prover was switched OFF for the runs (its sp1-gpu-server would otherwise be shared through `/tmp/sp1-cuda-0.sock` and carry its own environment; `gpu_server_before running=0`) and ON again at the end. Fixtures: four consecutive live blocks cut from PC 2's own node (86165..86168 at tip 86195, one empty shard each, every one MATCHES natively), the same four for every phase of job 1. App 0.3.9 on PC 2 throughout.
Known-finished case of the host changes before the GPU (this Mac, CPU, run lock, 20:41Z to 20:44Z): `--mode chain` over `fixtures/chain/block-81046.json` with `--save-shards` (shard 38.5 s, aggregate 43.4 s, the proof file written), then `--mode aggregate` over that saved shard proof with `SP1_WORKER_VERIFY_INTERMEDIATES=false` (46.6 s, the same statement `0x3dedb8ea...`), `--mode verify-segment` VERIFIED in 0.027 s; known-failed: a wrong statement NOT VERIFIED in 0.027 s. Unit tests: `cargo test --release -p igneum-prove-core -p igneum-prove-host`: core 8 passed, host 9 passed and 1 ignored (build lock, 20:53Z).
@ -2681,3 +2681,103 @@ Against the 5090 on the same PC (122 MH/s at 308 W, 0.396 MH/W): 25.3 percent of
Consequences per tier (the rule of 5 October 2026): a 5060 Ti owner (16 GB, Windows) mines at 30.9 MH/s and 115 W from the box with nothing to set: about 5,100 blocks a day at the 522 MH/s the devnet showed at 14:44Z (one every 17 s, approximate: the network rate moves), about a quarter of a 5090 owner's 20,200, for 2.76 kWh a day (£0.79 at 28.5 p against the 5090's £2.11); through a Thunderbolt enclosure the x4 link costs nothing measurable (the hash is bound by the card's own memory latency, not the link; the 5090's PCIe-slot rows are the comparison), so a laptop with a Thunderbolt 4 port and this enclosure is a 31 MH/s miner. The 8 GB 5060 Ti: the same hash is the expectation (the 1 GiB dataset fits), a line owed. Proving on the 16 GB tier: the fleet's 4060 Ti 16 GB row (9.0 GB peak beside the miner on the patched server) says this card would mine and prove with about 7 GB spare, approximate until the host with the device selector ships; today the app's prover default leaves it off ("a full shard needs a 24 GB card") and the measured read is owed to the prover-floor host cut. Linux and HiveOS take the same CUDA worker (owed a line). What the lane does next: the prover-floor host's selector into the shipped WSL2 bundle, then the prove-beside read on this card; the Power Helper fault on PC 2 to the Ember lane (no efficient point on any PC 2 card until it answers).
Found on the way: inside a PowerShell `@( ... )` the comma binds before `+`, so `'--query-gpu=' + $f, '--format=csv'` is one argument (run a, void in 1 s; the query string is built first now); a bare string inside a function that also returns a value is swallowed into the caller's variable (the sampler line; `[Console]::Out.WriteLine` now); the app's `kind` for a Thunderbolt card reads `discrete` (a word for the Cards page to earn: `external`, which the state already names).
## O-1.14: the CPU verifier on a 2019-class core (attack pass F6, 7 October 2026, 09:49 UK)
Rented Vast instance 54613164, Intel Core i7-9700K at 4,170 MHz as read, one core, the box-built Linux `igneum-pow`
(sha256 6d286783...), `bench --seed igneum-genesis --day 2026-10-03 --warps 50`. Ms per warp, cold max / average of 50:
v2 1.582 / 1.280; mx8 5.394 / 5.267; mx8+sh256x27 (class v4) 6.334 / 6.006; dr368 5.540 / 5.426; dr736 10.290 / 10.042
(fails the 10 ms gate, the known-fail). Cache fill 276 ms. Log `docs/analysis/attack-pass/o114-i7-9700K-2026-10-07.log`;
record `docs/analysis/attack-pass-2026-10.md` row F6. Class v4 passes a real 2019-class core with 3.7 ms to spare; the
half-core proxy (8.23 ms) stays the standing pessimistic rule for the ladder's ceiling.
## F6: the verifier's worst case over 10^5 class v4 programs (attack pass, 7 October 2026, 13:5x UTC)
igneum-build-1, cores 40 (one-core proxy) and 88 (the half-core proxy, both SMT siblings busy) under the per-core lease,
core 40 at a median 3,799.9 MHz. 100,000 programs ranked by exact op counts; 50,000 timed cold on core 40 (min / median /
p99 / max 4.610 / 4.948 / 5.606 / 6.194 ms per warp); the worst 200 re-timed at 10 cold reps on both proxies and the worst
1,000 on the half-core at 2 reps, the worst 10 at 20: worst half-core cold max 8.708 ms (`attack-f6/87142`), then 8.629
(`attack-f6/88521`), 8.414 (`attack-f6/15781`); the genesis program 8.624. Gate 10 ms: PASS by 1.29 ms. Record
`docs/analysis/attack-pass/f6-verifier.md`; logs `/srv/builds/igneum-wt-attack/target-attack-f6/phase2b.log`, `phase2c.log`.
## 7 to 8 October 2026, the class v4 efficiency passes: the core clock lock on the RTX 5090 and the RTX 5080 (branch ca3-v4-amend, the hash lane)
Main's question of 7 October ("profitable mining is all about efficiency, is there a solution?"): the latency-shadow work hides in the memory wait, so the core clock, and the voltage along the driver's V/F curve, can drop until the work just fits the wait, with no rate loss and the class v4 premium coming back. Measured on PC 1 (machine ae432dc7, driver 617.14, app 0.3.20) with `tools/ca3-v4-amend/pc1-v4-efficiency.ps1`: the installed CUDA worker's `--bench` on the class v4 pack `v4-devnet-epoch0` (id a785001687d8688a) and the class v3 control `mx8-devnet-epoch0` of the same seed (the two differ by the shadow block alone), the card alone through the runner's card switch, 60 s per step (the batch count sized from a 5-batch probe), the clock lock through the installed app's Power Helper task (`nvidia-smi -lgc 0,<MHz>`, a cap; no prompt), nvidia-smi at 1 Hz with the mean from 8 s in (power.draw, power.draw.instant and power.draw.average agree within 0.2 W on every row on 617.14), the 2^24 fingerprint at base 0 checked against the Mac on every step (every step matched), the clocks reset and read back at the end. Jobs: run-ca3-pc1-v4-eff-5090-20261007-b (18:40 to 19:16 UTC, 2,850 to 1,400 MHz), run-ca3-pc1-v4-eff-5090-floor-20261007 (20:21 to 20:41 UTC, 1,400 to 1,100; the steps below were lost to the helper's sequence rule, see the fud ledger), run-ca3-pc1-v4-eff-5080-20261007-d (8 October, 00:45 to 01:54 UTC, unlocked to 1,000 MHz plus 900 on v4; the 75 min budget ended it there). The memory clock stayed at the driver's default throughout (13,801 MHz on the 5090, 14,801 on the 5080).
RTX 5090 (lock MHz: class v4 MH/s / W / MH/W ; class v3 MH/s / W / MH/W; the SM clock read):
| lock | v4 MH/s | v4 W | v4 MH/W | v3 MH/s | v3 W | v3 MH/W | sm MHz |
|---|---|---|---|---|---|---|---|
| unlocked | 136.84 | 475.5 | 0.288 | 136.59 | 330.2 | 0.414 | 2,838 / 2,850 |
| 2850 | 136.89 | 478.9 | 0.286 | 136.61 | 330.3 | 0.414 | 2,833 / 2,842 |
| 2781 | 136.87 | 457.6 | 0.299 | 136.62 | 317.9 | 0.430 | 2,767 |
| 2700 | 136.85 | 443.2 | 0.309 | 136.54 | 312.4 | 0.437 | 2,692 |
| 2550 | 136.77 | 402.9 | 0.339 | 136.43 | 288.5 | 0.473 | 2,542 |
| 2472 | 136.62 | 391.2 | 0.349 | 136.38 | 276.0 | 0.494 | 2,460 |
| 2400 | 136.67 | 382.8 | 0.357 | 136.29 | 269.8 | 0.505 | 2,392 |
| 2250 | 136.43 | 366.5 | 0.372 | 136.07 | 255.0 | 0.534 | 2,242 |
| 2163 | 136.33 | 361.3 | 0.377 | 136.11 | 252.4 | 0.539 | 2,152 |
| 2100 | 136.25 | 359.3 | 0.379 | 136.04 | 251.4 | 0.541 | 2,092 |
| 1950 | 135.99 | 350.1 | 0.388 | 135.78 | 244.6 | 0.555 | 1,942 |
| 1854 | 135.85 | 341.4 | 0.398 | 135.56 | 239.6 | 0.566 | 1,845 |
| 1800 | 135.75 | 337.1 | 0.403 | 135.48 | 237.4 | 0.571 | 1,792 |
| 1650 | 135.37 | 328.1 | 0.413 | 135.10 | 235.1 | 0.575 | 1,642 |
| 1500 | 135.22 | 320.5 | 0.422 | 134.90 | 232.1 | 0.581 | 1,492 |
| 1400 | 134.98 | 316.3 | 0.427 | 134.68 | 228.0 | 0.591 | 1,387 |
| 1300 | 134.76 | 312.5 | 0.431 | 134.62 | 223.3 | 0.603 | 1,290 |
| 1200 | 133.80 | 305.1 | 0.439 | 129.54 | 215.7 | 0.601 | 1,192 |
| 1100 | 122.43 | 287.3 | 0.426 | 118.70 | 209.4 | 0.567 | 1,087 |
Reading (5090): the rate is memory-bound on the whole grid and holds within 1.5 percent of unlocked down to 1,300 MHz on both classes; it falls past 5 percent at 1,200 on class v3 and past 10 percent at 1,100 on class v4, so the knee is 1,300 MHz (42 percent of the 3,090 MHz maximum). Best MH per watt: class v4 at 1,200 MHz (133.80 MH/s at 305.1 W, 0.439 MH/W; 168.6 W recovered for 2.2 percent rate), class v3 at 1,300 MHz (134.62 at 223.3 W, 0.603; 106.6 W for 1.4 percent). The class v4 premium over class v3 is 145.3 W at the unlocked clock and 81.8 W at the best points: 63 W of the premium is the clock and comes back with the lock, 82 W is the shadow's ALU work and stays. The throttle reason reads the power governor (0x400) from unlocked to 2,163 and the clock lock itself (0x4) below. Consequence per tier: a 5090 owner on class v4 locked at 1,200 to 1,300 MHz draws 305 to 313 W instead of 474 for 1.5 to 2.2 percent less rate (MH per watt up 49 to 52 percent).
RTX 5080 (the same shape; 8 October 2026, 00:45 to 01:54 UTC):
| lock | v4 MH/s | v4 W | v4 MH/W | v3 MH/s | v3 W | v3 MH/W | sm MHz |
|---|---|---|---|---|---|---|---|
| unlocked | 71.41 | 253.1 | 0.282 | 71.28 | 169.7 | 0.420 | 2,963 / 2,977 |
| 2850 | 71.41 | 229.5 | 0.311 | 71.29 | 155.3 | 0.459 | 2,842 |
| 2781 | 71.41 | 223.9 | 0.319 | 71.29 | 150.1 | 0.475 | 2,767 |
| 2700 | 71.41 | 209.6 | 0.341 | 71.29 | 149.2 | 0.478 | 2,692 |
| 2550 | 71.41 | 193.0 | 0.370 | 71.29 | 134.4 | 0.531 | 2,542 |
| 2472 | 71.41 | 182.2 | 0.392 | 71.28 | 133.7 | 0.533 | 2,460 |
| 2400 | 71.41 | 176.0 | 0.406 | 71.29 | 128.7 | 0.554 | 2,392 |
| 2250 | 71.41 | 165.6 | 0.431 | 71.28 | 117.3 | 0.608 | 2,242 |
| 2163 | 71.39 | 161.5 | 0.442 | 71.28 | 114.1 | 0.625 | 2,147 / 2,152 |
| 2100 | 71.38 | 157.7 | 0.453 | 71.28 | 112.3 | 0.635 | 2,085 |
| 1950 | 71.37 | 154.0 | 0.463 | 71.26 | 111.6 | 0.639 | 1,942 |
| 1854 | 71.37 | 153.0 | 0.467 | 71.25 | 111.4 | 0.640 | 1,845 |
| 1800 | 71.35 | 150.3 | 0.475 | 71.24 | 113.4 | 0.628 | 1,777 / 1,785 |
| 1650 | 71.33 | 151.4 | 0.471 | 71.22 | 109.7 | 0.649 | 1,635 / 1,642 |
| 1500 | 71.30 | 149.2 | 0.478 | 71.19 | 110.6 | 0.644 | 1,492 |
| 1400 | 71.27 | 149.6 | 0.476 | 71.16 | 107.0 | 0.665 | 1,387 |
| 1300 | 71.24 | 147.9 | 0.482 | 71.14 | 107.7 | 0.661 | 1,275 / 1,282 |
| 1200 | 71.20 | 149.4 | 0.477 | 71.12 | 104.4 | 0.681 | 1,192 |
| 1100 | 71.20 | 146.6 | 0.486 | 71.11 | 105.6 | 0.673 | 1,087 |
| 1000 | 71.19 | 149.0 | 0.478 | 71.11 | 103.7 | 0.686 | 990 |
| 900 | 67.66 | 137.8 | 0.491 | not taken (the budget) | | | 893 |
Reading (5080): the rate holds within 0.3 percent of unlocked down to 1,000 MHz on both classes and falls 5.2 percent at 900 MHz on class v4, so the knee is between 1,000 and 900 MHz, a third of the 2,963 MHz boost and lower than the 5090's (the 5080's 84 SMs have more compute headroom per unit of its memory bandwidth, so the memory wait hides the shadow down to a lower clock). Best MH per watt within the 1 percent rate tolerance: class v4 at 1,100 MHz (71.20 MH/s at 146.6 W, 0.486 MH/W; 106.5 W recovered for 0.29 percent rate), class v3 at 1,000 MHz (71.11 at 103.7 W, 0.686; 66.0 W for 0.25 percent). The class v4 premium is 83.4 W unlocked and 41 W at the best points (146.6 W against 105.6 W at 1,100). The draw floors from 1,500 MHz down (about 147 W on v4, 104 W on v3) with the power governor as the only throttle reason say the clock lever is spent by 1,500 MHz on this card. Consequence per tier: a 5080 owner on class v4 locked near 1,100 MHz pays 147 W instead of 253 for 0.3 percent less rate (MH per watt up 72 percent).
What the lever is and is not: nvidia-smi exposes no voltage offset (that is NVAPI's); the clock lock walks the driver's V/F curve, which is where the watts come from; the Mac has no lever (no clock cap on Apple silicon); AMD has the ADLX tune line through igneum-gpu-telemetry or nothing. The knob goes into Ember Tune for 0.3.24 (src/ember.rs: the clock ladder continues below 45 percent of the maximum in 100 MHz steps to a 20 percent floor, the search stops at the first row more than the tolerance under the cap point's rate or on a faulted row, the best MH per watt within tolerance is the point, the fingerprint checked on every step, the result stored per card as the lock_* fields). The 5080 stock rows against the rented 5080 of 7 October (71.16 MH/s at 145 W on driver 580): the rate agrees to 0.4 percent, the watts do not (253 W here); PC 1's three power fields agree, so the difference sits with the rented card's sampler or its cap, the fleet lane's re-measure owed.
## 8 October 2026, the hot-table packs on the RTX 5090: ld.global.cs against the plain load, unlocked and at the 1,300 MHz lock (branch ca3-v4-amend, the hash lane, job run-ca4-pc1-hot-ldcs-5090-20261008b)
PC 1, RTX 5090 (170 SMs, driver 13.4, NVRTC 12.8), igneum-worker-cuda 1.0 (4 October 2026). The eight 5 October hot-table packs (proto-cuda/packs-ca2-hot) and the mx8-genesis control (proto-cuda/packs-ca2-mixer), each run twice per state: variant base (plain loads) and variant ldcs (ld.global.cs on the hot-table loads). 16,777,216 nonces per run, nvidia-smi 1 Hz sampler (power.draw; instant and average agreed within 0.5 W on every row), 19 to 30 samples per row. Every one of the 36 rows PASS on its pinned fingerprint. Unlocked 06:34 to 06:45 UTC at 2,865 MHz; lock1300 06:45 to 06:56 UTC at 1,290 MHz through the helper; clocks reset (rgc) at the end. The 5090 was switched off in the app by the runner for the job and back on after.
| pack | unlocked base MH/s | unlocked base W | unlocked ldcs MH/s | lock1300 base MH/s | lock1300 base W | lock1300 base MH/W | lock1300 ldcs MH/s | lock1300 ldcs MH/W |
|---|---|---|---|---|---|---|---|---|
| mx8-genesis (control) | 137.7 | 312.0 | 137.7 | 127.3 | 211.4 | 0.602 | 127.3 | 0.603 |
| hot32k4 | 147.4 | 318.9 | 147.4 | 144.4 | 215.3 | 0.671 | 144.4 | 0.672 |
| hot32k4a | 118.9 | 320.1 | 118.9 | 116.6 | 214.9 | 0.542 | 116.6 | 0.542 |
| hot64k2 | 137.7 | 318.1 | 137.7 | 135.0 | 214.3 | 0.630 | 135.0 | 0.630 |
| hot64k4 | 141.0 | 318.8 | 141.0 | 138.2 | 214.4 | 0.645 | 138.2 | 0.644 |
| hot64k4a | 115.7 | 319.6 | 115.8 | 113.4 | 214.8 | 0.528 | 113.4 | 0.528 |
| hot64k8 | 164.2 | 326.4 | 164.2 | 160.9 | 219.2 | 0.734 | 160.8 | 0.733 |
| hot96k4 | 139.0 | 318.3 | 139.0 | 136.2 | 214.5 | 0.635 | 136.3 | 0.635 |
| hot96k4a | 114.6 | 318.8 | 114.5 | 112.3 | 214.8 | 0.523 | 112.3 | 0.523 |
What the rows say.
1. The ldcs variant changes nothing: every pack reads the same rate and the same watts as its base run within 0.1 MH/s and 1 W, both states. The streaming hint on the hot-table loads is dead as a lever; the table's cache behaviour is already what the hardware gives. No further ldcs rows are owed.
2. The hot packs hold their rate under the lock far better than mx8: the 1,300 MHz lock costs mx8 7.5 percent of rate (137.7 to 127.3) and costs the hot packs 2 percent (hot32k4 147.4 to 144.4, hot64k8 164.2 to 160.9). The hot packs are bound by the table's latency, not by the core; the lock takes a third of the watts off every pack (319 to 215 W) and the hot packs pay almost no rate for it.
3. Per watt at the lock the hot family runs 0.52 to 0.73 MH/W against the control's 0.60: hot64k8 is 22 percent cheaper per hash than mx8 on this card, hot32k4 11 percent cheaper, the "a" packs 10 to 13 percent dearer. Whether a cheaper hash on the GPU is a gain or a loss for resistance is the research lane's call: it is a gain only if the saving comes from the memory path an ASIC would have to buy too.
4. The 5090's locked class v4 reading from the efficiency pass (1,300 MHz: 134.6 MH/s at 223 W, 0.60 MH/W) sits level with the mx8 control here (0.602), so the two passes agree on the control and the hot rows are comparable to the v4 grid.

View file

@ -204,7 +204,7 @@ key of 8.1 is the operator's own step, on the file.
| PC 2's 5090 rows | the card-off step could not confirm a stopped worker (`api/state` answered `{}`), and the numbers are half the card's | the morning's 5090-only run on PC 2 with the compute-apps confirmation; until then the rows are "beside the live miner" |
| PC 2's result files | `repro.ps1`'s close threw `System.OutOfMemoryException` in ConvertTo-Json and gave Set-Content an empty path: PowerShell variable names are case-insensitive, the result table `$Cpu` held `$cpu` (its own name string) and so contained itself, and the markdown lines `$md` wiped the markdown path `$Md`. Fixed (distinct names; `bench/jobs/ps-case-check.sh` fails CI on any case-only pair, shown to fire on a bad file); the RESULT lines in the intake and the three transcript logs are the record | the morning's run writes the files |
| The prover on PC 2 | the first script read the prover's state from `api/state` too, saw nothing, and left the live prover off from 00:49 to the restore job (run-prover-on-pc2-20261006); the script now reads settings.json and restores unconditionally | done tonight; the rule in the job |
| The reward | the amount and payer are `docs/plans/funding.md`, not funded; the terms are quoted unchanged | the project lead's decision |
| The reward | the amount and payer are `docs/plans/funding.md`, not funded; the terms are quoted unchanged | the founder's decision |
## 8. The vendor-share metric (Counter ASIC 3.0 item 7)
@ -221,7 +221,7 @@ Today's devnet, read-only from the live tables (dry mode, 07:36 UTC, 6 October 2
| Vendor | fleet-reported MH/s (workers) | chain-attributed share of blue blocks (10 min) | chain-attributed MH/s at the 135.8 MH/s estimate | Note |
|---|---|---|---|---|
| NVIDIA | 120.61 (1: PC 2's RTX 5090, with the prover on the card) | 0.986 (552 of 560) | 133.86 | PC 1's RTX 5090 was off the network at the reading (the project lead's desk) |
| NVIDIA | 120.61 (1: PC 2's RTX 5090, with the prover on the card) | 0.986 (552 of 560) | 133.86 | PC 1's RTX 5090 was off the network at the reading (the founder's desk) |
| AMD | 0 (0) | 0 | 0 | the RX 9070 XT is on PC 1, off at the reading |
| Apple | 0 (0) | 0 | 0 | the M5 Max is paused for measurements |
| Intel | 1.69 (1: the Windows laptop's UHD Graphics) | 0.014 (8 of 560) | 1.94 | the first outside machine, which carries the intake key |

View file

@ -0,0 +1,122 @@
# The light-client bridge primitive (Devnet 3 to Ethereum Sepolia), 8 October 2026
Devnet 3, test tokens, no value. This is a design plus one working contract. It moves nothing. It proves two things on
Sepolia and states, below, what it does not prove.
## What exists
| Piece | Where | What it does |
|---|---|---|
| The verifier contract | `contracts/bridge/src/IgneumCertificateVerifier.sol`, deployed on Sepolia (address in `docs/contracts/sepolia.json`) | holds a Devnet 3 voter table; verifies a finality certificate against it and records the checkpoint hash as final; verifies an Ethereum account proof against a state root |
| BLS12-381 | `contracts/bridge/src/BLS12381.sol` | hash-to-curve for G2 (RFC 9380, the node's DST), key aggregation in G1, the two-pairing check, all through the EIP-2537 precompiles |
| The account proof | `contracts/bridge/src/MerklePatricia.sol` | walks an eth_getProof-shaped proof (RLP nodes, hex-prefix paths, embedded children) to the account's RLP value or to its absence |
| The suite | `contracts/bridge/test/Verifier.t.sol`, vectors from `test/vectors/gen.mjs` | 9 tests on Foundry's Prague EVM: the vote message byte for byte, RFC 9380 expand_message_xmd answers, a certificate by four of five made-up keys, one under two thirds refused, forged index and checkpoint refused, the recorded checkpoint, an account present and an account absent under one root, a wrong root refused, and one certificate the chain actually carried |
| The vectors | `gen.mjs synthetic`, `gen.mjs table <weights.json>`, `gen.mjs chain <checkpoint.json>` | made-up keys and a three-account trie; the Devnet 3 voter table from a node's `igneum_getFinalityWeights`; a real certificate as `/api/checkpoint` serves it, keys and signature decompressed to the precompiles' encodings |
The suite ran green on build-3 (9 of 9) before the deploy; the Sepolia address and the deploy transactions are in
`docs/contracts/sepolia.json`.
## What the certificate check proves
A certificate is `(index, checkpoint hash, bitmap, aggregate signature)` as `consensus/core/src/finality.rs` writes it
into a block's coinbase (`Certificate::write`: index_le64, checkpoint, voter_count_le32, bitmap_len_le32, bitmap,
signature, aggregator, aggregator proof). The contract takes the first four fields; the signature is the 96-byte G2
point decompressed off chain to the precompiles' 256-byte form (a wrong decompression is a point the pairing
precompile refuses).
The contract holds a voter table installed by its deployer: the canonical voter list at one checkpoint index (every
key above dust and not stripped, sorted by key hash, as the node's `igneum_getFinalityWeights` reports it), each key a
128-byte uncompressed G1 point, each weight the key's blue blocks in the 30-day window. `tableId` is the keccak of
`(index, keys, weights)`.
`verifyCertificate(index, checkpoint, bitmap, signature)` holds exactly when:
1. bit `p` of the bitmap names voter `p` of the table, and the sum of the named voters' weights is at least two thirds
of the table's total weight (the rule decided 4 October 2026: lock = 2/3 of all 30-day weight; stricter than the
17/30 floor plus 2/3 of active that the node line still carries, so every certificate the node locks under the new
rule passes here and some the old rule locked would not);
2. the aggregate of the named keys (G1 additions) verifies the signature over the vote message
`"igneum-vote-v1/" || chain_id || 0x00 || index_le64 || checkpoint` under the domain separation tag
`IGNEUM_VOTE_V1_BLS12381G2_XMD:SHA-256_SSWU_RO_NUL_`, with `chain_id` the string the contract was built with
(`igneum-devnet-3`): e(aggregate key, H(message)) * e(-G1, signature) = 1.
So a recorded `finalCheckpoint(index)` means: the keys in the installed table that hold at least two thirds of that
table's weight signed this checkpoint hash at this index on this chain id. That is the finality rule's own statement,
checked with the node's own bytes, by a contract on another chain.
Measured on the test EVM: a 29-voter table costs about 5.3 million gas to install; a certificate with 12 signers
verifies in about 3.9 million gas (the pairing and the two map-to-curve calls dominate; a G1 addition per signer is
375 gas). On Sepolia at a 1 gwei tip that is under 0.01 ETH per certificate.
## What the account proof proves
`verifyAccount(stateRoot, account, proof)` walks an Ethereum account proof (the `eth_getProof` shape: RLP nodes from
the root, keys hashed with keccak, the hex-prefix leaf and extension paths, children by hash or embedded when under 32
bytes) and returns the account's nonce, balance, storage root and code hash, or `exists = false` when the trie shows the
account absent. This is the trie Igneum's executor commits to: `igneum/exec/src/state.rs` computes `stateRoot` with
`alloy_trie::root::state_root` over `keccak256(address)` and `RLP(nonce, balance, storage_root, code_hash)`, EIP-161
empty accounts left out, reth's layout, the same as Ethereum's.
So a verified account proof means: under this state root, this account has these fields.
## What is NOT proven, in order of weight
1. **The link from a certified checkpoint to a state root.** Nothing the contract verifies ties a checkpoint hash to an
EVM state root. The Kaspa-shaped header the voters sign has no execution root (its `utxo_commitment` is the UTXO
multiset, `accepted_id_merkle_root` the accepted transaction ids); the executor runs behind consensus as a follower
and the EVM block hash is the DAG block hash, not a hash of the EVM header. The binding that exists today is in the
coinbase of later blocks: the IGNS segment record (`BlockStatement.post_root` for its last chain block, signed by
the aggregator's vote key) and the proof records (`ProofRecord.statement` = keccak of a `ShardOutput` carrying
`pre_root` and `post_root`, signed by the prover's vote key), both under the carrier block's `hash_merkle_root`. A
carried record that fails a node's native veto is ignored rather than faulting the block, so even that binding
rests on the aggregator's and prover's keys and on the nodes' native check, not on a header field. Today a caller
gives the verifier a state root as a stated input beside a verified certificate; the contract does not know they
belong together. The reference-apps lane's oracle verifies the coinbase path (BLAKE2b header hash, the merkle
branch, the record parse) on Sepolia and shares this verifier for the certificate; the two together are the full
chain once the record's signature check lands there.
2. **The voter table itself.** The table is installed by the deployer from a node's `igneum_getFinalityWeights` read:
the keys, their canonical order (sorted by `BLAKE2b-256("IgneumVoteKeyHash", key)`, which the contract does not
recompute: no BLAKE2b on chain today) and their weights are trusted. A wrong table makes a wrong verdict in both
directions. A light client proper tracks the table from the chain: every weight is a count of blue blocks whose
headers name the key, so the table follows from headers; and the node's `FinalityWeights` RPC reports the table at
the certificate's own index (`voters_at_index`). The next step is a table update that takes a certified checkpoint
plus the headers between locks, the shape `/api/checkpoint` already serves to the browser verifier
(`site/verify/core.js` checks exactly that path, in JavaScript).
3. **The checkpoint index is a number the submitter gives.** The contract records one hash per index and refuses a
second, but it does not know the chain's current index; an old certificate at an old index verifies for ever
against the table it was signed under. A consumer should read `finalCheckpoint` at an index it already knows from
the chain, not treat the newest recorded index as the chain's tip.
4. **Equivocation and stripping.** The node strips a key's weight for 30 days on equivocation evidence. The installed
table carries the stripping as of its read and nothing after it.
5. **Proof of work, the DAG order, execution correctness.** None of it is checked here. The certificate's claim is the
voters' signature, and the voters are the miners who proved their blocks; the ZK proofs of execution (the shard
proofs the records name) are verified by the nodes, not by this contract. A proof-carrying bridge (the chain's SP1
shard proofs verified on Ethereum) is the design's end state and is not today's primitive.
6. **The state root's age and the account's current balance.** An account proof says what the balance was under that
root; the root is one block's. Nothing here prevents a stale root from being presented.
## How one balance gets proven on Sepolia today, and what each step rests on
| Step | Source | Rests on |
|---|---|---|
| the voter table at index N | a Devnet 3 node's `igneum_getFinalityWeights` | the node, the installer (trusted today) |
| the certificate for checkpoint C at index N | `/api/checkpoint?source=dn3` (the observer's `dn3_live_certificates`, the bytes a block carried) | verified on chain: the signature and the two-thirds rule |
| the state root R of chain block B | `eth_getBlockByNumber` on Devnet 3 | stated, not proven against C (gap 1) |
| the account proof for A under R | `eth_getProof` on the reference-apps lane's reader node (fork branch light-apps-node, the Devnet 3 pin plus a read-only RPC) | verified on chain against R |
The test `test_account_proof_present_and_absent` proves the account under a made-up root; the Devnet 3 account
proof against a real root goes into the suite as `test/vectors/dn3-account.json` the moment the reader node serves
`eth_getProof` (the proof shape is Ethereum's, so the contract needs no change).
## The design from here
1. The header path on chain: BLAKE2b-256 through the EIP-152 precompile for the keyed header hash and the key hashes,
the merkle branch to the coinbase, the segment and proof records parsed from the coinbase payload. This closes gap 1
and lets the contract recompute canonical order (gap 2's order).
2. The table update from certified headers: weights counted from the blue blocks between two locks, so the table
follows the chain instead of an installer. This closes gap 2.
3. A tip rule: the contract keeps the highest index it has recorded and a consumer reads only at or below it; the
relayer submits each lock as it lands (one transaction per 30-second checkpoint is affordable on Sepolia, not on
mainnet; mainnet gets one certificate per epoch).
4. The proof-carrying bridge: the shard proofs' aggregate verified on Ethereum (the prover network's own product),
which replaces trust in the native veto with a verified execution root.

View file

@ -7,13 +7,13 @@ shard run reported as exit 0, 7a7e873).
| Date | Symptom | Cause | Fix | Proven by |
|---|---|---|---|---|
| 4 Oct 2026 | Every `ci` run on master red since 67bf226 (eleven pushes), unnoticed | `sim/difficulty/records/testnet-v2-2026-10-04.schedule.log` carried a home path; `.log` was outside the identity scrub's extension list in `tools/ci/identity-check.sh` (and in the mirror's `tools/sync.sh`) | 2996cca: `.log` scrubbed like the other text files; the record rewritten with `~`; the same list in igneum-public `tools/sync.sh` (local commit e18256d, not pushed) | `bash tools/ci/identity-check.sh` 0 hits locally; run 37226816xxx on master green |
| 5 Oct 2026 | PC 1 (Windows 11 Pro 26200, default terminal Windows Terminal 1.24): "Windows Command Processor" windows whenever a remote job runs (the project lead) | measured, not guessed: `tools/windows/console-watch.ps1` (job run-20261005-182528) started every candidate child from the app's job runner, whose console is headless (`conhost.exe 0x4`, hwnd 0), with a user32 EnumWindows sampler every 30 ms: powershell, cmd, query, curl, nvidia-smi, wsl --status, a distro, interop cmd and powershell, `powershell -WindowStyle Hidden`, `Start-Process -WindowStyle Hidden`: 0 windows each; `Start-Process cmd` in a new console: a Terminal window and a cmd PseudoConsoleWindow (the known-failed case fires). The 25-minute background watcher (console-watch-bg.ps1, run-20261005-184330, 18:44 to 19:09 UTC, every 200 ms) across an app restart, a build job, two run jobs, two collect jobs and the sweep helper's elevated launch at 19:04:43: 0 console or Terminal windows, 69 conhost starts (every one `conhost.exe 0x4`, headless, under curl, wsl, wslhost, powershell), 1 cmd.exe (under wslhost, WSL interop, no window). The one road that creates a console of its own is the elevated launch (`Start-Process -Verb RunAs`, the AppInfo service: the power cap, the sweep helper, the clock sync, an elevated job); it carried `-WindowStyle Hidden` in four copies, and "Windows Command Processor" is also the name on the UAC prompt the engine raises for cmd.exe (the sweep helper prompted at 17:00, 17:30 and 18:12 UTC, the power cap at every start; the elevated watcher's own prompt, run-20261005-184610, timed out unanswered at 122 s) | `platform::elevated_ps_line` + `elevated_command`: one builder for every elevated launch, hidden by construction, exit 251 when the prompt is refused; the elevated job wrapper reports its own console (`elevated console: hwnd N visible False`) on every elevated job; `tools/ci/windows-spawn-check.mjs` fails CI on a Command::new without the quiet flag, a creation_flags other than CREATE_NO_WINDOW, a Start-Process without -WindowStyle Hidden/-NoNewWindow, or a host.cpp spawn without CREATE_NO_WINDOW / SW_HIDE | the watcher's known-failed case (2 windows) and known-finished case (0); the CI check's self-test (9 cases) and the tree (0 hits); the igneum-app test suite on PC 1 |
| 5 Oct 2026 | PC 1 (Windows 11 Pro 26200, default terminal Windows Terminal 1.24): "Windows Command Processor" windows whenever a remote job runs (the founder) | measured, not guessed: `tools/windows/console-watch.ps1` (job run-20261005-182528) started every candidate child from the app's job runner, whose console is headless (`conhost.exe 0x4`, hwnd 0), with a user32 EnumWindows sampler every 30 ms: powershell, cmd, query, curl, nvidia-smi, wsl --status, a distro, interop cmd and powershell, `powershell -WindowStyle Hidden`, `Start-Process -WindowStyle Hidden`: 0 windows each; `Start-Process cmd` in a new console: a Terminal window and a cmd PseudoConsoleWindow (the known-failed case fires). The 25-minute background watcher (console-watch-bg.ps1, run-20261005-184330, 18:44 to 19:09 UTC, every 200 ms) across an app restart, a build job, two run jobs, two collect jobs and the sweep helper's elevated launch at 19:04:43: 0 console or Terminal windows, 69 conhost starts (every one `conhost.exe 0x4`, headless, under curl, wsl, wslhost, powershell), 1 cmd.exe (under wslhost, WSL interop, no window). The one road that creates a console of its own is the elevated launch (`Start-Process -Verb RunAs`, the AppInfo service: the power cap, the sweep helper, the clock sync, an elevated job); it carried `-WindowStyle Hidden` in four copies, and "Windows Command Processor" is also the name on the UAC prompt the engine raises for cmd.exe (the sweep helper prompted at 17:00, 17:30 and 18:12 UTC, the power cap at every start; the elevated watcher's own prompt, run-20261005-184610, timed out unanswered at 122 s) | `platform::elevated_ps_line` + `elevated_command`: one builder for every elevated launch, hidden by construction, exit 251 when the prompt is refused; the elevated job wrapper reports its own console (`elevated console: hwnd N visible False`) on every elevated job; `tools/ci/windows-spawn-check.mjs` fails CI on a Command::new without the quiet flag, a creation_flags other than CREATE_NO_WINDOW, a Start-Process without -WindowStyle Hidden/-NoNewWindow, or a host.cpp spawn without CREATE_NO_WINDOW / SW_HIDE | the watcher's known-failed case (2 windows) and known-finished case (0); the CI check's self-test (9 cases) and the tree (0 hits); the igneum-app test suite on PC 1 |
| 4 Oct 2026 | `collect-pc1-board3` printed PowerShell parse errors (`.Name`, `.AdapterRAM`) | the publishing shell expanded `$_` inside double quotes to nothing before the command reached the jobs file; nothing to do with Format-List or Out-String (board2 and board4 printed their values) | publish-jobs.sh refuses a collect command that pipes into a script block without `$_` or `$PSItem` | the eaten form refused with the reason, the single-quoted form published to a test folder |
| 4 Oct 2026 | the same job reported `done (exit 0)` over `command exit Some(1)` | `run_collect` in `app/igneum-app/src/jobrun.rs` builds `Done` from the upload count only; the command's exit code is logged and dropped | branch `bugfix-collect-exit`, 35ccdc8 rebased on c257444 (app engine; merge by the main session) | `cargo test --bin igneum-app`: all 28 tests pass on the rebased branch; the new one covers the board3 shape (`Some(1)` is failed exit 1), `Some(0)` done, the cap as timeout, failed uploads still failing |
| 4 Oct 2026 | `publish-jobs.sh --deploy` said "not reachable, differs from the local one, or does not verify yet" after a deploy that had succeeded | one check the instant the CLI returned, while the edge still served the previous file; the deploy's own exit status was hidden by `\|\| true` | `verify_live`: up to `--tries` (12) checks 5 s apart, each failure names its condition; `publish-jobs.sh verify` re-checks on its own; a failed deploy stops before the check | finished: `verify --tries 2` against the live file (try 1 of 2); failed: a local server with an older file ("differs", both publish stamps named) and a closed port ("is not reachable") |
| 4 Oct 2026 | console Machines: PC 37ba0461 showed 0.0 MH/s and 0 accepted while its log held an accepted block at 1 MH/s | `parseLabel` in the console API knew nvidia, amd, mac, metal and opencl; the OpenCL fallback on an iGPU is labelled `other-<id8>-n` and the card was dropped | parsers moved to `relay/lib/parse.mjs`, vendors `other` and `intel` added, `node --test relay/test/parse.test.mjs` in CI | the test; the live console after the deploy shows the card |
| 4 Oct 2026 | console Machines: a card said "117.2 MH/s now" while its "status" column said 4 m ago (PC 2 during shard run 3: the prover held the GPU and the worker's STATUS line stopped) | the card's hash came from the last STATUS line in the tail with no age check; the machine total summed it | `markStale`: a card whose STATUS line is older than 120 s is `stale`, shown as "last N MH/s" with a red "stale" mark, and left out of the machine total (API, page and `tools/console.mjs`) | the test (58 s fresh, 240 s stale, none stale); the live console after the deploy |
| 4 Oct 2026 | `vercel env add` from `site/` fails with "Could not retrieve Project Settings" | `site/.vercel/project.json` links the [other-business] team's `igneum` project; the live site (igneum.network, igneum.com, the GitHub integration) is the `igneum` team's project of the same name, which the igneum login reads and the [other-business] link does not | documented in `packaging/README-ship.md` (link, env ls, env add, deploy); no env set | `env ls` from a scratch link to the igneum-team project lists the two names; a branch push produced `igneum-git-<branch>` |
| 4 Oct 2026 | `vercel env add` from `site/` fails with "Could not retrieve Project Settings" | `site/.vercel/project.json` links the other business's team's `igneum` project; the live site (igneum.network, igneum.com, the GitHub integration) is the `igneum` team's project of the same name, which the igneum login reads and the the other business link does not | documented in `packaging/README-ship.md` (link, env ls, env add, deploy); no env set | `env ls` from a scratch link to the igneum-team project lists the two names; a branch push produced `igneum-git-<branch>` |
| 4 Oct 2026 | `tools/jobs.mjs` and `publish-jobs.sh` print the downloads-folder token inside URLs on every run (X24 pattern) | the tokened base URL is echoed as is | the token masked as `<token>` in every printed URL | by eye, this log's own transcript |
| 4 Oct 2026 | round 4 X28 and X24, the parts under an hour: `===` on secrets, no HSTS on the relay, the relay token printed by `tools/relay.mjs list` and `watch` | as the review said | c1f59fb: `sameSecret` (timingSafeEqual, `relay/lib/auth.mjs`, test in CI), `Strict-Transport-Security` in `relay/vercel.json`, `/r/<token>` printed (only `url` prints the real one) | relay deployed: key auth 200, wrong key and token 401, token path 200, HSTS header present |
| 4 Oct 2026 | the live feed showed two "checkpoint N locked" events 30 ms apart (1122, 1172, 1230, 1258, 1259 in 400 observer lines), and 708 of 764 locked checkpoints in `live_checkpoints` had `votes_seen` 0 | `finalityTick` read the state, awaited two SQL writes, then set the map; the `finalityLockNotification` handler checked the same map synchronously in between and recorded the lock too; a lock claimed by the notification was never upserted again, so the poll's `votes_seen` never landed | 7de1bdb: the poll claims the state before its first await; a `checkpointDetailed` set makes the poll fill `votes_seen` once | before: 5 duplicates in 400 lines; after the 19:50:33 UTC restart: 21 locks (1297 to 1317), 0 duplicates, 0 write failures; 1300 was notification-first and the poll filled it to 17 votes. The zeros that remain are the node's own count (`finality.rs:770`, its vote map for that hash, empty when the lock came by certificate), not the observer's. Index 1296 appears twice on the feed: it locked inside the restart window, one write per process, a restart-boundary one-off At the 20:15:58 restart index 1319 (locked 14 min before) was recorded again: open, the seed should have held it locked; f22870a logs the seeded states and the earlier state on such a record. The 20:24:46 restart seeded 'proposed 500, locked 864' and re-recorded nothing (11 locks, 0 duplicates) |

16
docs/build/README.md vendored Normal file
View file

@ -0,0 +1,16 @@
# docs/build: the builder programme
8 October 2026. These files are the source of the developer pages on igneum.network; `site/build.mjs` renders two of them as pages and the rest are read here.
| File | Serves |
|---|---|
| `build.md` | [igneum.network/build](https://igneum.network/build): what is different (built to prove every block, a lock in minutes, nothing to stake, the EVM unchanged), the chain ids and endpoints, the wallet, the explorer, the faucet, the five-minute contract, the browser verifier |
| `grants.md` | [igneum.network/grants](https://igneum.network/grants): the tiers, what a grant is, how to apply, the review, the honest lines |
| `first-contract.md`, `first-contract-test.sh` | the walkthrough and the script that runs it end to end on Devnet 3 from a build box; the PASS line is the test record |
| `rpc.md` | the JSON-RPC method list as the node serves it, read from a Devnet 3 node, with the probe |
| `faucet.md` | the Devnet 3 faucet: rules, where it runs, how it is funded |
| `verify-a-block.md` | the in-browser checkpoint verifier, twelve lines |
The faucet service is `infra/build-server/faucet/`. The grant application template is `.forgejo/issue_template/grant.md`.
Served-text rules (the site audit reads every page before it lands): no em dashes, short sentences, one to three points per block, nothing that reads as an offer, a price or a promise of value; devnet and testnet coins have no value; no amount, price or date is promised.

88
docs/build/build.md vendored Normal file
View file

@ -0,0 +1,88 @@
# Build on Igneum
This file is the source of [igneum.network/build](https://igneum.network/build). The site build renders it as the page. The other files in this directory are the walkthrough, the RPC list, the faucet, the verifier and the grants.
## What is different
**Every block is built to be proven.** A chain block carries a zero-knowledge proof of its execution. The miners are the provers: the same cards that find blocks prove them, in shards. On Devnet 3 a share of blocks carries a proof today; the live page shows the share and the lag, and the design target at launch is under a minute. Full nodes execute every block themselves, so a bad proof is a light-client problem and never a chain split.
**A lock in minutes, not an hour of confirmations.** Miners sign a checkpoint every 30 seconds. When two thirds of the mining weight of the last 30 days have signed, the checkpoint is locked. Weight is blocks mined, nothing else. The lock lands in about two minutes. The [litepaper](/litepaper) has the rule.
**Nothing to stake.** No validator set, no delegation, no slashing, no bonded class above the miners. Finality comes from mining weight alone. Nobody holds a key that the chain depends on.
**The EVM you already know.** Solidity deploys unchanged. Standard JSON-RPC, EIP-1559 transactions, the chain id in the signature, `cancun` as the EVM version. Gas has two dimensions on Igneum, execution and proving, and the node folds the second into the price it quotes, so `eth_estimateGas` and `eth_gasPrice` work as they do on Ethereum.
## Networks
| | Devnet 3 | Testnet | Mainnet |
|---|---|---|---|
| Chain id | 4463 (`0x116f`) | 4462 (`0x116e`) | 4461 (`0x116d`) |
| Network id | `igneum-devnet-3` | `igneum-testnet-1` | not started |
| RPC | `https://rpc.devnet.igneum.network` (a node you run serves `http://127.0.0.1:26790`) | `https://rpc.testnet.igneum.network` (answers; nothing mines there yet, so a transaction waits) | none |
| Coins | no value, resets without notice | no value, resets with notice | not started |
| Symbol, decimals | IGN, 18 | IGN, 18 | IGN, 18 |
Devnet 3 is where you build today. It is a developer network: it resets without notice and its coins have no value. Its public RPC takes the read methods, `eth_sendRawTransaction` and a wRPC websocket at `/ws`, at 20 requests a second per address. The public testnet exists, its RPC answers, and no blocks are being produced on it until it opens. Mainnet has no date. Devnet 3 answers chain id 4463 until its class v5 floor, then 4464.
## Endpoints and tools
- **Wallet.** Any Ethereum wallet. [Add Igneum to MetaMask](/metamask) with one click, or the [Igneum wallet](/wallet).
- **Faucet.** [10 IGN of Devnet 3 coin](/faucet) per address per day. No account.
- **Explorer.** [The explorer](/explorer) shows blocks, the DAG, miners and addresses. The EVM views, transactions and contracts, are being built and link from the same page as they land.
- **Swap.** [The swap](/swap): test tokens on Devnet 3, every swap in a proven block.
- **Source.** [git.igneum.network/igneum-network/igneum](https://git.igneum.network/igneum-network/igneum): the node, the miner, the site, and these pages under `docs/build/`.
## Your first contract in five minutes
With [Foundry](https://getfoundry.sh). Tested end to end on Devnet 3 on a build box; the full walkthrough with the expected output is [docs/build/first-contract.md](https://git.igneum.network/igneum-network/igneum/src/branch/master/docs/build/first-contract.md).
```
curl -L https://foundry.paradigm.xyz | bash && foundryup
export RPC=https://rpc.devnet.igneum.network
cast wallet new # an address and a private key, for the devnet only
```
Paste the address into [the faucet](/faucet). Then:
```
export PK=0x... # the private key cast printed
cast balance $(cast wallet address $PK) --rpc-url $RPC --ether
forge init counter && cd counter
forge create src/Counter.sol:Counter --rpc-url $RPC --private-key $PK --broadcast
```
`forge create` prints `Deployed to: 0x...`. Call it:
```
cast send 0xDEPLOYED "increment()" --rpc-url $RPC --private-key $PK
cast call 0xDEPLOYED "number()(uint256)" --rpc-url $RPC
```
The call answers `1`. Paste the transaction hash into [the explorer](/explorer). Hardhat, ethers and viem work the same way: chain id 4463, the RPC above.
## The RPC the node serves
Read from a Devnet 3 node, not from a spec. `web3_clientVersion` answers `igneumd/2.1.0/execution-layer-v3`. The full list, served and not served, with the probe that produced it: [docs/build/rpc.md](https://git.igneum.network/igneum-network/igneum/src/branch/master/docs/build/rpc.md).
- **Served.** `eth_chainId`, `eth_blockNumber`, `eth_getBalance`, `eth_getTransactionCount`, `eth_getCode`, `eth_getStorageAt`, `eth_getBlockByNumber`, `eth_getBlockByHash`, `eth_getTransactionByHash`, `eth_getTransactionReceipt`, `eth_getBlockReceipts`, `eth_getLogs`, `eth_call`, `eth_estimateGas`, `eth_gasPrice`, `eth_maxPriorityFeePerGas`, `eth_feeHistory`, `eth_sendRawTransaction`, `eth_syncing`, `eth_accounts`, `net_version`, `net_peerCount`, `net_listening`, `web3_clientVersion`.
- **Igneum's own.** `igneum_getExecStatus`, `igneum_getProvingStatus`, `igneum_getShardPlan`, `igneum_getProofRecords`, `igneum_getTransactionStatus`, `igneum_estimateGas`, `igneum_getBudgets` and the segment and proof calls. `igneum_estimateGas` returns both gas dimensions and whether a call exceeds the proving limit.
- **Not served.** Filters and subscriptions (`eth_newFilter`, `eth_subscribe`), `txpool_*`, `debug_*`, `trace_*`, `eth_sendTransaction`, `eth_sign`, `eth_getProof`, the uncle calls. Poll `eth_blockNumber` and `eth_getLogs`. Sign on your side.
## Verify a block in your browser
The site ships a light client. Your browser fetches the latest certified checkpoint from `/api/checkpoint` and verifies the BLS aggregate signature, the voter weights and the header chain itself, in the tab. Nothing is trusted from the server but the data. The example, in twelve lines, is [docs/build/verify-a-block.md](https://git.igneum.network/igneum-network/igneum/src/branch/master/docs/build/verify-a-block.md); the live run is [verify/test.html](/verify/test.html), genuine and tampered cases side by side.
```
import { fetchCheckpoint, verify } from 'https://igneum.network/verify/verify.js';
const data = await fetchCheckpoint('https://igneum.network/api/checkpoint');
const r = verify(data);
console.log(r.verified, r.signers, 'of', r.voters, 'voters', r.ms, 'ms');
```
## Where to ask
- An issue on [the repository](https://git.igneum.network/igneum-network/igneum/issues).
- [Discord](https://discord.gg/igneum), the builders channel.
- [Grants](/grants) for tooling, reference apps, infrastructure and research.
Devnet and testnet coins have no value. Nothing on this page is an offer to sell anything.

36
docs/build/faucet.md vendored Normal file
View file

@ -0,0 +1,36 @@
# The Devnet 3 faucet
[igneum.network/faucet](https://igneum.network/faucet) sends 10 IGN of Devnet 3 coin to an address. Once per address per day, once per connection per day. No account.
## What it is
- A small Node service on a build box, `infra/build-server/faucet/faucet.mjs`, behind Caddy at `faucet.igneum.network`. The page on igneum.network posts to it.
- It signs a plain transfer with its own key and sends it to the Devnet 3 node on the same box by address (`FAUCET_RPC`), never chosen by chain id: Devnet 3 answers the same chain id as the shared devnet.
- The key lives only in the service's environment file on that box. It is not in the repository and not on the site.
## The rules, as the code enforces them
| Rule | Value |
|---|---|
| Amount | 10 IGN per send; the code refuses to start with an amount over the 100 IGN cap |
| Per address | one send per 24 hours |
| Per connection | one send per 24 hours |
| Per day, in total | 500 sends, then "come back tomorrow" |
| What is stored | the address, the connection's IP, the transaction hash and the time, in a file on the box, for the 24-hour window |
The faucet's balance is refilled from the devnet's funder as it runs down. When it is empty the page says so and sends nothing.
## Devnet 3 coin
No value. Not for sale, not redeemable, not a claim on anything. The chain resets without notice and balances do not carry over. Mainnet starts from an empty genesis.
## Running it
```
infra/build-server/faucet/faucet.sh install # dry run: what would be copied and started
infra/build-server/faucet/faucet.sh install --go # copy the service, the unit and the Caddy block; start it
infra/build-server/faucet/faucet.sh status # the unit, the balance, the sends in the last 24 hours
node infra/build-server/faucet/faucet.mjs --self-test # the limiter: known-failed cases first
```
The public testnet faucet is a separate thing: `site/api/faucet.mjs` is prepared for it and answers "not open yet" until the testnet starts.

49
docs/build/first-contract-test.sh vendored Executable file
View file

@ -0,0 +1,49 @@
#!/usr/bin/env bash
# The five-minute walkthrough (docs/build/first-contract.md), run end to end on a build box against a Devnet 3 node, so the
# page is tested and not described. Runs where Foundry is installed (~/.foundry/bin); asks the faucet for the coins exactly
# as a reader would, deploys Foundry's Counter, calls it, reads the receipt, prints one PASS line with the numbers.
# RPC=http://127.0.0.1:27810 FAUCET=http://127.0.0.1:8790/api/faucet bash docs/build/first-contract-test.sh
# Known-failed first: with RPC pointed at a closed port the script must print FAIL and exit 1 (--self-test).
set -uo pipefail
export PATH="$HOME/.foundry/bin:$PATH"
RPC="${RPC:-https://rpc.devnet.igneum.network}"
FAUCET="${FAUCET:-https://faucet.igneum.network/api/faucet}"
t0=$(date +%s); now() { date -u +%H:%M:%SZ; }
fail() { echo "FAIL $(now): $*"; exit 1; }
if [ "${1:-}" = --self-test ]; then
out=$(RPC=http://127.0.0.1:9 FAUCET=http://127.0.0.1:9/x bash "$0" 2>&1); rc=$?
[ "$rc" = 1 ] && case "$out" in *FAIL*) echo "self-test passed: a closed RPC port reads FAIL, exit 1"; exit 0 ;; esac
echo "self-test failed: rc=$rc out=$out"; exit 1
fi
command -v forge >/dev/null || fail "forge is not installed"
cid=$(cast chain-id --rpc-url "$RPC" 2>/dev/null) || fail "no chain id from $RPC"
[ "$cid" = 4463 ] || [ "$cid" = 4464 ] || fail "chain id $cid is not Devnet 3"
work=$(mktemp -d); cd "$work" || fail "no scratch dir"
# step 1: a key for the devnet only (never printed)
w=$(cast wallet new 2>&1 | sed 's/\x1b\[[0-9;]*m//g') || fail "cast wallet new"
PK=$(printf '%s\n' "$w" | sed -n 's/^Private key: *//p' | head -1); ME=$(printf '%s\n' "$w" | sed -n 's/^Address: *//p' | head -1)
[ -n "$PK" ] && [ -n "$ME" ] || fail "cast wallet new printed no key"
# step 2: the faucet, as a reader would
f=$(curl -s -m 60 -X POST -H 'content-type: application/json' --data "{\"address\":\"$ME\"}" "$FAUCET") || fail "the faucet did not answer"
case "$f" in *'"ok":true'*) ;; *) fail "the faucet refused: $f" ;; esac
drip=$(printf '%s' "$f" | sed -n 's/.*"tx":"\(0x[0-9a-f]*\)".*/\1/p')
for i in $(seq 1 60); do bal=$(cast balance "$ME" --rpc-url "$RPC" --ether 2>/dev/null); case "$bal" in 10.*|9.*) break ;; esac; sleep 1; done
case "$bal" in 10.*|9.*) ;; *) fail "balance after the faucet: ${bal:-none}" ;; esac
# step 3: deploy Foundry's template counter
forge init counter --no-git >/dev/null 2>&1 || forge init counter >/dev/null 2>&1 || fail "forge init"
cd counter || fail "no counter dir"
out=$(forge create src/Counter.sol:Counter --rpc-url "$RPC" --private-key "$PK" --broadcast 2>&1) || fail "forge create: $(printf '%s' "$out" | tail -3)"
C=$(printf '%s\n' "$out" | sed -n 's/^Deployed to: *//p' | head -1); deployTx=$(printf '%s\n' "$out" | sed -n 's/^Transaction hash: *//p' | head -1)
[ -n "$C" ] || fail "no Deployed to line: $out"
# step 4: call it
cast send "$C" "increment()" --rpc-url "$RPC" --private-key "$PK" >/dev/null 2>&1 || fail "cast send increment"
n=$(cast call "$C" "number()(uint256)" --rpc-url "$RPC" 2>/dev/null); [ "$n" = 1 ] || fail "number() after increment: $n"
sendOut=$(cast send "$C" "setNumber(uint256)" 41 --rpc-url "$RPC" --private-key "$PK" 2>&1) || fail "cast send setNumber"
txh=$(printf '%s\n' "$sendOut" | sed -n 's/^transactionHash *//p' | head -1)
n=$(cast call "$C" "number()(uint256)" --rpc-url "$RPC" 2>/dev/null); [ "$n" = 41 ] || fail "number() after setNumber: $n"
# step 5: the receipt
st=$(cast receipt "${txh:-$deployTx}" --rpc-url "$RPC" 2>/dev/null | sed -n 's/^status *//p' | head -1)
case "$st" in 1*|*success*) ;; *) fail "receipt status: $st" ;; esac
blk=$(cast block-number --rpc-url "$RPC" 2>/dev/null)
echo "PASS $(now): Devnet 3 chain id $cid, faucet drip $drip, deployed Counter at $C (tx $deployTx), increment then setNumber(41) read back 41, receipt status $st, block $blk, $(( $(date +%s) - t0 )) s end to end, Foundry $(forge --version 2>/dev/null | sed -n 's/^forge Version: *//p' | head -1)"
cd / && rm -rf "$work"

Some files were not shown because too many files have changed in this diff Show more