From 2f2bdb5a705ebd7d1aeed5ca8bb0049ee7b1949f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:42:12 +0000 Subject: [PATCH 1/3] plan: the prover-floor agent's first sweep (the v1 shard at 12.7 GB with the 2^26 split, 5.3 s; the 9.7 GB Setup is next) and the 12 GB profile decision Co-Authored-By: Claude Fable 5.1 --- docs/plans/proving-v1.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index b7ac3175..b2fa92e6 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -114,6 +114,8 @@ Reading. Nobody pays an aggregator as a separate role: Aztec's 30% goes to whoev The resume path (5 October 2026, the 0.3.11 app): `POST /api/resume` on 0.3.9 re-armed only FAULTED cards (`stop_miners("paused")` clears every slot's `restart_at`), so a healthy paused card stayed "off" at 0 MH/s until the app was relaunched: PC 2 at 21:25:11Z (the aggregation-cost job's pause and resume; `[ok] mining resumed` then `0.00 MH/s, waiting` for 20 minutes), the Mac that afternoon. Now every slot without a live worker is re-armed and its pack exported again before the start, and 90 s later `resume_check` logs `resume: is not mining 90 s after resume (state ..., pid ...)` for every enabled card without a hash rate (`engine.rs`, three unit tests: the state machine, the 21:25:11Z case against the old rule, the check). +The prover-floor agent's first sweep (job `floor-sweep-1`, 22:34 to 22:38Z, PC 2's 5090, the miners stopped, this plan's per-point recipe, its patched `sp1-gpu-server` 5568108b built for sm_86, sm_89 and sm_120, every proof VERIFIED by the unpatched pv1 host): the control at upstream's sizes reproduces the curve above (empty shard 13,892 MiB and 2.2 s; the v1 shard 20,516 MiB and 4.2 s); with the core element threshold at 2^26 the v1 shard proves as four core shards in 5.3 s at **12,708 MiB** and the empty shard at 12,772 MiB; 2^25 gives 12,836 MiB at 8.5 s; 2^27 gives 15,396 MiB. The 12.7 GB left is the server's Setup (five recursion keys pre-built at a fixed 2^27 capacity plus the shrink and core keys: 9.7 GB before the first shard), which its patch v2 sizes to the need. Decided for the 12 GB profile: the split that lands under 11 GB wins (5.3 s a shard is inside the loop's own 25 to 30 s of carriage and 100x inside T); 2^27 is the second profile only if v2 leaves it under 11 GB with the miner's 1.8 GB beside it. The 12 GB row stays OPEN until the final pair (alone and beside the miner) lands and the on-order 3060 runs it. + ### A self-built CUDA server (the 12 GB path), before 0.3.12 (consequences C26) If the prover-floor agent's rebuilt `sp1-gpu-server` (the Setup sizes cut, built on PC 2 under WSL2) proves a shard under 11 GB, it becomes a shipped artefact and needs its own row of rules before 0.3.12: it is built from a pinned SP1 source tag with `CUDA_ARCHS` covering sm_86, sm_89 and sm_120 (the 12 and 16 GB tiers are Ampere and Ada, not only the 5090's Blackwell; one card family per measured row), by the packaging path that builds the Windows payload (PC 1's build job for the Linux binary, the Mac signs the manifest as it does the DMG), lands in the DMG and the WSL2 package beside the host as `wsl2/bin/sp1-gpu-server` with its sha256 in `payload-inputs.json`, is named in `evidence.md` beside the prover rows ("prover built from SP1 at "), is rebuilt and re-measured at every SP1 upgrade, and ships only after `--mode verify-segment` and `--mode verify` on proofs it made show the pinned verifying keys unchanged (the server changes allocation, not the circuit; the ids `0x2b1a81cb...` and `0x474678f3...` must still verify them). The 12 GB claim itself waits for the on-order RTX 3060 to run that server on the same fixtures and recipe as the curve; until then the public line stays at 24 GB. From 047932f2e14f80499f7b1dce78244d64afed1322 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 00:18:32 +0000 Subject: [PATCH 2/3] plan: the prover-floor agent's sweep 3 rows (the patched server proves the v1 shard at 10.3 GB on the 12 GB profile, 12.9 GB at 2^27; the 12 and 16 GB rows move from nothing to proves alone, the pairs pending) Co-Authored-By: Claude Fable 5.1 --- docs/plans/proving-v1.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index b2fa92e6..8badea48 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -132,8 +132,8 @@ The GPU server of SP1 6.8.1 sets the memory, not the shard: a floor of 13.9 GB f |---|---|---|---| | 32 GB (RTX 5090) | the prototype shard, 28.3 GB, 10.8 s; the v1 shard 20.4 GB, 4.3 s | the prototype shard 30.1 GB, 33 s; the v1 shard 22.2 GB, 13.2 s | on, mine and prove, today | | 24 GB (RTX 4090, 3090) | the v1 shard 20.4 GB; the prototype shard does NOT fit (28.3 GB) | the v1 shard 22.2 GB measured on the 5090's allocation (2.3 GB spare on a 24 GB card; approximate for the card itself) | on, mine and prove, with the line "until the devnet's fee switch its shards are the prototype size, which needs 32 GB, so this card proves from the switch on" | -| 16 GB (RTX 5080, 4080) | an empty shard only (13.9 GB) | nothing (15.7 GB for an empty shard, no room for the display) | off, with the line | -| 12 GB (RTX 3060, 4070) | nothing: the floor is 13.9 GB, and the shipped server refuses the card outright | nothing | off; the project lead's "make sure we can prove on 12 GB cards" is OPEN and in work: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (`sp1-gpu/crates/prover_components/src/builder.rs` lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) | +| 16 GB (RTX 5080, 4080) | the shipped server: an empty shard only (13.9 GB); the patched server v3 b37defef at threshold 2^27: the v1 shard 12,915 MiB and 4.3 s, the prototype shard 13,459 MiB and 16.8 s (measured by the prover-floor agent on the 5090's allocation, job `floor-sweep-3`, 00:13 to 00:17Z 6 October; 2^27 + 2^26 gives 16,115 MiB, over the card) | the shipped server: nothing (15.7 GB for an empty shard); the patched server: about 14.7 GB at 2^27 beside the miner (approximate: the measured 1.8 GB the miner adds; the pair is the agent's sweep 4) | off on the shipped server, with the line; on once the patched server ships (the packaging row below) and the pair is measured | +| 12 GB (RTX 3060, 4070) | the shipped server: nothing (the floor is 13.9 GB, and the server refuses the card outright); the patched server v3 b37defef at threshold 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`, the 12 GB profile): **the v1 shard 10,291 MiB and 5.7 s, an empty shard 9,971 MiB and 3.3 s**, the card's 2,089 MiB idle inside the peak and the server's own working set about 8.2 GB (6,535 MiB after Setup), so a 12 GB card proves alone with about 3 GB over it (measured by the prover-floor agent on the 5090's allocation, `floor-sweep-3`; the on-order RTX 3060 run is pending) | the pair (the miner's 1.7 to 1.8 GB and 3x beside it) is the agent's sweep 4; under 9.0 GB mine-and-prove is not yet shown | off on the shipped server; "proves alone" on the patched one once it ships (the packaging row below), mine-and-prove after the pair. the project lead's "make sure we can prove on 12 GB cards" is answered on the 5090's allocation and OPEN on the card itself: the prover-floor agent (branch prover-floor, 5 October night) read SP1 v6.8.1's GPU server source (`sp1-gpu/crates/prover_components/src/builder.rs` lines 35 to 39): it reads the card's memory, adds 4 and panics under 24 ("Unsupported GPU memory ... must be at least 24GB"), and builds its core (ELEMENT_THRESHOLD 2^28 + 2^27 elements + 2^21), recursion (2^27), shrink (2^25) and wrap (85 M element) provers at Setup whatever the mode, which is the 13.9 GB floor; no knob reaches them, so the fix is a server rebuilt from source on PC 2 (WSL2, nvcc 12.8, CUDA_ARCHS=120) with those sizes cut, measured on the same fixtures and recipe as the curve above (D2 carries the curve) | | under 12 GB | nothing | nothing | off, mine only | | AMD-only and Apple machines | nothing on the GPU: no zkVM proves on an AMD GPU today (`docs/analysis/amd-proving.md`, branch amd-prove); the CPU prover is about 5 minutes a shard at a 30 GB RSS whatever the shard size (PC 1, bench-log "the SP1 CPU prover on PC 1") | | off, "mines and does not prove"; the only non-NVIDIA path with a shipped backend is RISC Zero's Metal prover behind the `ProofSystem` seam (a second guest and pinned id, a verifier for both formats, no shared aggregation): an open item, not 0.3.11 | From 42f36b3fd3ad21d7d91562c6d000f66acef7063b Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 00:27:35 +0000 Subject: [PATCH 3/3] app: /api/state never answers {} again (paid_wei over u64::MAX broke to_value; the field is a decimal string, the error is logged once) serde_json's to_value refuses a u128 over u64::MAX (18.45 IGN) and state_json turned the error into json!({}). A paid shard is 1.23 IGN on average, so a proving machine's dashboard went blank about 15 paid shards after every app start. Unit test over the boundary; 114 app tests pass. Co-Authored-By: Claude Fable 5.1 --- app/igneum-app/src/engine.rs | 15 ++++++++++++++- app/igneum-app/src/state.rs | 22 ++++++++++++++++++++++ docs/bench-log.md | 15 +++++++++++++++ docs/plans/proving-v1.md | 7 +++++++ 4 files changed, 58 insertions(+), 1 deletion(-) diff --git a/app/igneum-app/src/engine.rs b/app/igneum-app/src/engine.rs index 5278f089..d67acb7e 100644 --- a/app/igneum-app/src/engine.rs +++ b/app/igneum-app/src/engine.rs @@ -92,6 +92,8 @@ pub struct Shared { /// the app log's path: the one the uploader sends (main.rs names it; the engine must not guess the stamp) pub log_path: PathBuf, port: std::sync::atomic::AtomicU16, + /// set once `state_json` has logged a serialisation error (one line, not one per poll) + state_error_logged: std::sync::atomic::AtomicBool, } impl Shared { @@ -135,6 +137,7 @@ impl Shared { engine_log: Mutex::new(engine_log), log_path, port: std::sync::atomic::AtomicU16::new(0), + state_error_logged: std::sync::atomic::AtomicBool::new(false), } } @@ -177,7 +180,17 @@ impl Shared { st.now = crate::platform::unix_now_f(); st.uptime_s = self.started.elapsed().as_secs(); st.events = self.rings.lock().unwrap().events.iter().cloned().collect(); - serde_json::to_value(st).unwrap_or(json!({})) + match serde_json::to_value(st) { + Ok(v) => v, + Err(e) => { + // never a silent "{}": the dashboard and every PC playbook read this reply + let msg = format!("state_json: {e}"); + if !self.state_error_logged.swap(true, std::sync::atomic::Ordering::Relaxed) { + self.log(&format!("[error] {msg}")); + } + json!({ "error": msg, "version": env!("CARGO_PKG_VERSION") }) + } + } } /// A short state for the window host (menu bar, tray). diff --git a/app/igneum-app/src/state.rs b/app/igneum-app/src/state.rs index 5cf32e57..da328308 100644 --- a/app/igneum-app/src/state.rs +++ b/app/igneum-app/src/state.rs @@ -172,6 +172,9 @@ pub struct ProvingState { pub submitted: u32, pub paid: u32, pub failed: u32, + /// wei, serialised as a decimal string: serde_json's `to_value` refuses a u128 over u64::MAX (about 18.45 IGN, + /// 15 paid shards at 1.23 IGN), and that refusal emptied the whole `/api/state` reply to "{}" (6 October 2026) + #[serde(serialize_with = "u128_string")] pub paid_wei: u128, pub current: String, pub started_at: f64, @@ -401,3 +404,22 @@ impl Rings { out } } + +/// A u128 as a decimal JSON string (the dashboard reads it with `Number()`). +pub fn u128_string(v: &u128, s: S) -> Result { + s.serialize_str(&v.to_string()) +} + +#[cfg(test)] +mod paid_wei_tests { + use super::*; + + #[test] + fn a_paid_total_over_u64_max_still_serialises_the_whole_state() { + let mut st = State::default(); + st.proving.paid_wei = u64::MAX as u128 + 1; + let v = serde_json::to_value(&st).expect("the state serialises"); + assert_eq!(v["proving"]["paid_wei"], serde_json::Value::String("18446744073709551616".into())); + assert!(v["mining"].is_object()); + } +} diff --git a/docs/bench-log.md b/docs/bench-log.md index 2976bd59..a0ebfa9e 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1694,3 +1694,18 @@ Reading, with the miner-on pairs above (empty shard 15,670 MiB, full prototype s |---|---| | Unit tests | `cargo test --release -p kaspa-consensus-core -p igneum-exec --lib -- proving config::params::tests::override_params_carry_the_proving_v1 config::params::tests::consensus_digest` on this Mac (target `vendor/igneum-node/target-pv1`, 19:09Z): consensus core 13 passed (the segment record round trip, signature and the three nested sections; the credit split; the params switch and the digest that moves only once the switch is set), exec 8 passed (the segment grid and the split; the record checks: alignment, block, chain length, the veto naming the field, the deadline, the window; the chain rule both ways; the unproven restart; the shard side at 90%; the pool offering the segment section). The six full node suites go to PC 2 as a build job when the fleet is back | | The fast-time 3-node harness (`tools/proving-v1/net.mjs`, 29950+, suffix 956, every node in trust mode, three vmine voters, v0 at DAA 60, v1 at DAA 120, 4 blocks a segment, unproven after 60 DAA, a tenth to the aggregator; fork b177718e built on this Mac) | run 2, 19:13:01Z to 19:16:19Z, under the run lock: PASSED, 21 checks in 197.3 s (`tools/proving-v1/report-2026-10-05.json`). v1 start = chain block 119 on all three nodes; the native statement identical on all three. Known-finished: segment 119..122's fresh-chain record submitted to n1 at t=131.1 s, relayed, verified (trust) and PAID on n0 1.0 s later at chain block 129, 253,611,648,000,000,000 wei = a tenth of the four credits, the same on every node, the payout address holding it. Chain rule: segment 123..126's fresh-chain record refused ("does not chain to segment 119..122 ... proven (record paid at chain block 129)"), the continuing one (chain_len 8) accepted and paid. Known-failed: segment 127..130 left without a record: a fresh-chain record for 131..134 refused while 127..130 was pending ("pending until DAA 191"); at DAA 192 the status read unproven, a late record for 127..130 refused ("unproven: carried after the deadline"), the fresh-chain record for 131..134 accepted and paid with chain_len 4; `segmentsInWindow` proven 3, unproven 1. The shard side: a v1 shard's `shardWei` = 90% of its block's credit. Run 1 (19:10Z) failed in its own tooling (the signer's argument order), fixed. Run 3 on the FINAL fork tree (ece42979 on the 0.3.10 commit 21d4c73c, protocol 15, N = 8 both in the params default and `--segment 8`, the fast-time file's four fields), 20:52:41Z to 20:56:45Z: PASSED, 21 checks in 244.4 s (segments of 8: 119..126 paid in 1.0 s after submission, 127..134 refused fresh and paid continuing with chain_len 16, 135..142 left unproven and skipped, 143..150 restarted the chain) | + +### 6 October 2026, 00:4xZ, the empty `/api/state` reply (proving v1 branch) + +Reported by the aggregation-cost agent: PC 2's `/api/state` answered `{}` (2 bytes) at 22:22Z, 22:41Z and 00:18Z. Not measured on PC 2 (no job); derived from the app source and node 1's RPC, read-only on the Mac: + +| Figure | Value | Source | +|---|---|---| +| Paid shards, devnet, all provers | 663 | `curl -s 127.0.0.1:26790 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}'` at tip DAA 0x22caf | +| Paid wei, all provers | 0x2c2961a69990745400 = 814.64 IGN | same call | +| Average per paid shard | 1.23 IGN (approximate: the mean over 663) | 814.64 / 663 | +| u64::MAX in IGN | 18.45 | 2^64 - 1 over 1e18 | +| Paid shards per app start before the reply empties | 15 (approximate: at the mean payout) | 18.45 / 1.23 | + +Cause: `ProvingState.paid_wei: u128` and serde_json `to_value` (1.0.151, `value/ser.rs` `serialize_u128`: u64 range or an error); the error became `json!({})`. Fix: the field serialises as a decimal string; `state_json` logs the error once. Test `a_paid_total_over_u64_max_still_serialises_the_whole_state` (`cargo test --offline -q paid_wei`, 1 passed). + diff --git a/docs/plans/proving-v1.md b/docs/plans/proving-v1.md index 8badea48..e19e5959 100644 --- a/docs/plans/proving-v1.md +++ b/docs/plans/proving-v1.md @@ -143,3 +143,10 @@ The aggregation-cost agent's first rows (branch agg-cost, 5 October 2026 night, The re-plans of block 344 at 2.25 M and 4.5 M pgas peak at 28.3 to 28.4 GB alone (the server's buffers step up between 4.7 M and 20 M cycles and are flat to 60 M), so no shard size between the v1 budget and the prototype one changes a tier; with the miner the adopted shard proves 3.1x slower (13.2 s against 4.2 s) and the chained aggregation 9.7 s against 2.5 s: a mining 24 GB card delivers one adopted-size shard plus one aggregation in about 23 s, inside T by 25x. +## The empty `/api/state` reply (6 October 2026) + +The aggregation-cost agent's jobs read the two bytes `{}` from `/api/state` on PC 2 at 22:22Z, 22:41Z and 00:18Z (0.3.10 and 0.3.11); the 21:01Z reply was full. Cause, from the app source and node 1's RPC: `ProvingState.paid_wei` is a `u128`, and serde_json's `to_value` refuses a u128 over u64::MAX (18,446,744,073,709,551,615 wei, 18.45 IGN); `state_json()` turned that refusal into `json!({})` with no log line. A paid shard is 1.23 IGN on average (node 1, `igneum_getProvingStatus`: 814.64 IGN over 663 shards at 00:3xZ), so the fifteenth paid shard after an app start empties the reply. PC 2's prover was blind to the root-owned socket from 20:00:56Z to 22:01Z (paid_wei stayed 0, hence the full reply at 21:01Z), proved from 22:02:13Z, and crossed 18.45 IGN inside its first 15 paid shards, before 22:22Z. Every app restart resets the counter, so the reply comes back for about 15 shards and goes again. + +What it means: the dashboard on a proving machine shows nothing within about 12 minutes of its prover's first payout; every PC playbook that reads a card from `/api/state` fails the same way (the agent's job 5 reads settings.json instead). Mining, proving and payouts are untouched; it is the status page only. + +Fix on the app branch: `paid_wei` serialises as a decimal string (the dashboard already reads it with `Number()`), `state_json` logs `[error] state_json: ...` once instead of answering `{}`, and the reply on any future serialisation error carries `error` and `version` rather than nothing; unit test `a_paid_total_over_u64_max_still_serialises_the_whole_state`. Ships with 0.3.11 if the shipper takes the new code tip, otherwise 0.3.12; until then the workaround is settings.json for the card keys.