Relay: run and task posts need the console token (round 4, X23); prove host saves proofs buffered (ledger P20, second gap)
The intake key sits in every miner package, so the relay now lets it report only (drop text and files, ack, done, register, upload). Posting a run or task, or renaming and re-roling a machine, needs the console token. The prove host wrote proofs through SP1's unbuffered save: on WSL2 under /mnt/c the 18 MB core proof of a shard took longer to save than to prove. Proofs now go through a 4 MB buffer with a timed 'saved' line, and prove-shard.sh keeps results on the Linux side and copies them per stage. Ledger P20 updated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
860c1ec6ca
commit
42e6ec9c81
5 changed files with 33 additions and 7 deletions
|
|
@ -1472,7 +1472,7 @@ Evidence: `docs/bench-log.md`, 4 October 2026 "execution layer attack fixes" (be
|
|||
### P20. The SP1 GPU client panics on shutdown and the compressed stage waited ten minutes
|
||||
"Your first GPU proof run aborted with a core dump. What else aborts?"
|
||||
|
||||
Status: Fixed in the host, to be confirmed on the PC (4 October 2026, afternoon). (1) `igneum-prove-host` holds a Tokio runtime for the whole run (`main` enters it before anything else) and drops the proof system inside it, then shuts the runtime down with a 10-s grace, so `sp1-cuda`'s `tokio::spawn` in `Drop` finds a runtime. (2) Every stage prints a `STAGE ... start` line and a `RESULT` line with a UTC timestamp, and the setup line splits client creation from the two key setups (on the Mac CPU the client creation alone is 30 to 50 s of a 40 to 60 s setup; bench-log 4 October 2026, shards). The next PC run (PROVE-SHARD.bat) shows whether the gap sits before the first compressed stage and how long it is on the card. Was: Open, found by the team (4 October 2026, morning, first RTX 5090 run through WSL2, SP1 6.8.1).
|
||||
Status: (1) fixed in the host; (2) found on the second RTX 5090 run (4 October 2026, evening): the silence sits BEFORE the `STAGE compressed` line, so it is not the recursion setup. The only work there is saving the core proof, and SP1's `save` serialises straight into an unbuffered file: on WSL2 with the results folder under /mnt/c every 4-byte field element is one round trip across the Windows file bridge, and the 18.1 MB core proof of a full shard had not finished saving 25 minutes after the proof verified (the block-78 run's 10 minutes were the same save of a smaller proof). Fix: the host writes proofs through a 4 MB buffer and prints `saved <file> in N s`, and `prove-shard.sh` writes results on the Linux side and copies them to the Windows folder once per stage. Pending the timed line from the next PC run. Was: Fixed in the host, to be confirmed on the PC (4 October 2026, afternoon). (1) `igneum-prove-host` holds a Tokio runtime for the whole run (`main` enters it before anything else) and drops the proof system inside it, then shuts the runtime down with a 10-s grace, so `sp1-cuda`'s `tokio::spawn` in `Drop` finds a runtime. (2) Every stage prints a `STAGE ... start` line and a `RESULT` line with a UTC timestamp, and the setup line splits client creation from the two key setups (on the Mac CPU the client creation alone is 30 to 50 s of a 40 to 60 s setup; bench-log 4 October 2026, shards). The next PC run (PROVE-SHARD.bat) shows whether the gap sits before the first compressed stage and how long it is on the card. Was: Open, found by the team (4 October 2026, morning, first RTX 5090 run through WSL2, SP1 6.8.1).
|
||||
|
||||
Answer: Two separate things, neither in the proof. (1) After every proof was written, verified and uploaded, `sp1-cuda`'s client dropped its session key outside a Tokio runtime and panicked in its destructor (`sp1-cuda-6.8.1/src/pk.rs:63`, `client.rs:221`), so the host exited 134 with the results already on disk. Fix in our host: hold a runtime for the client's lifetime or drop the proof system inside one. (2) Between the core proof (08:49:10 UTC) and "Proving with mode: Compressed" (08:59:02 UTC) the host was silent for ten minutes while the card was idle; the compressed mode's recursion setup on first use is the suspect, and the second run must time it. Both go on the proving e2e benchmark standard as fixed overheads to measure, not hide.
|
||||
|
||||
|
|
|
|||
|
|
@ -13,6 +13,7 @@ alloy-primitives.workspace = true
|
|||
alloy-trie.workspace = true
|
||||
alloy-rlp.workspace = true
|
||||
bincode.workspace = true
|
||||
serde.workspace = true
|
||||
serde_json.workspace = true
|
||||
anyhow.workspace = true
|
||||
hex.workspace = true
|
||||
|
|
|
|||
|
|
@ -301,6 +301,24 @@ fn run_execute(sp1: &Sp1ProofSystem, shards: &[BuiltShard], parent_hash: B256, r
|
|||
results.insert("aggregator_cycles".into(), cycles.into());
|
||||
Ok(())
|
||||
}
|
||||
/// Writes a proof with a 4 MB buffer and prints how long it took. `SP1ProofWithPublicValues::save` serialises
|
||||
/// straight into an unbuffered `File`, which on WSL2 means millions of tiny writes across the /mnt/c bridge:
|
||||
/// the 18 MB core proof of a full shard took the RTX 5090 run over 15 minutes to save (ledger P20, second gap).
|
||||
fn save_proof<T: serde::Serialize>(proof: &T, path: std::path::PathBuf) {
|
||||
let t = Instant::now();
|
||||
let res = std::fs::File::create(&path)
|
||||
.map_err(|e| anyhow!("{e}"))
|
||||
.and_then(|f| {
|
||||
let mut w = std::io::BufWriter::with_capacity(4 << 20, f);
|
||||
bincode::serialize_into(&mut w, proof).map_err(|e| anyhow!("{e}"))?;
|
||||
std::io::Write::flush(&mut w).map_err(|e| anyhow!("{e}"))
|
||||
});
|
||||
match res {
|
||||
Ok(()) => println!("saved {} in {:.1} s", path.display(), t.elapsed().as_secs_f64()),
|
||||
Err(e) => println!("could not save {}: {e}", path.display()),
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
fn run_shard(sp1: &Sp1ProofSystem, shards: &[BuiltShard], index: usize, out_dir: Option<&str>, results: &mut serde_json::Map<String, serde_json::Value>) -> Result<()> {
|
||||
let s = shards.get(index).ok_or_else(|| anyhow!("shard {index} is not in the plan ({} shards)", shards.len()))?;
|
||||
|
|
@ -327,7 +345,7 @@ fn run_shard(sp1: &Sp1ProofSystem, shards: &[BuiltShard], index: usize, out_dir:
|
|||
results.insert("core_verify_seconds".into(), vdt.as_secs_f64().into());
|
||||
results.insert("core_proof_bytes".into(), bytes.into());
|
||||
if let Some(dir) = out_dir.and_then(|p| std::path::Path::new(p).parent()) {
|
||||
let _ = core.save(dir.join(format!("block-{}-shard-{i}-core.bin", s.input.env.number)));
|
||||
save_proof(&core, dir.join(format!("block-{}-shard-{i}-core.bin", s.input.env.number)));
|
||||
}
|
||||
|
||||
stage(&format!("compressed shard {i}"));
|
||||
|
|
@ -343,7 +361,7 @@ fn run_shard(sp1: &Sp1ProofSystem, shards: &[BuiltShard], index: usize, out_dir:
|
|||
results.insert("compressed_verify_seconds".into(), vdt.as_secs_f64().into());
|
||||
results.insert("compressed_proof_bytes".into(), bytes.into());
|
||||
if let Some(dir) = out_dir.and_then(|p| std::path::Path::new(p).parent()) {
|
||||
let _ = proof.proof.save(dir.join(format!("block-{}-shard-{i}-compressed.bin", s.input.env.number)));
|
||||
save_proof(&proof.proof, dir.join(format!("block-{}-shard-{i}-compressed.bin", s.input.env.number)));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
|
@ -470,7 +488,7 @@ fn run_block(sp1: &Sp1ProofSystem, shards: &[BuiltShard], claim: &SegmentClaim,
|
|||
results.insert("block_proof_bytes".into(), bytes.into());
|
||||
results.insert("block_total_seconds".into(), total.into());
|
||||
if let Some(dir) = out_dir.and_then(|p| std::path::Path::new(p).parent()) {
|
||||
let _ = seg.proof.save(dir.join(format!("block-{}-aggregated.bin", seg.output.number)));
|
||||
save_proof(&seg.proof, dir.join(format!("block-{}-aggregated.bin", seg.output.number)));
|
||||
}
|
||||
let _ = BlockOutput::LEN;
|
||||
Ok(())
|
||||
|
|
|
|||
|
|
@ -46,17 +46,19 @@ if ! cargo build --release -p igneum-prove-host --features igneum-prove-host/cud
|
|||
fi
|
||||
HOST="$DEST/proving/igneum-prove/target/release/igneum-prove-host"
|
||||
FIXDIR="$DEST/proving/fixtures"
|
||||
mkdir -p "$HERE/results"
|
||||
mkdir -p "$HERE/results" "$DEST/results" # proofs are written on the Linux side (ext4) and copied to the Windows folder once per stage: every byte through /mnt/c costs a bridge round trip
|
||||
|
||||
echo "=== GPU shard run: SP1_PROVER=cuda, $SHARD_FIXTURE, shard 0: execute + core + compressed (the first run downloads sp1-gpu-server, about 134 MB) ==="
|
||||
SP1_PROVER=cuda RUST_LOG=info "$HOST" "$FIXDIR/$SHARD_FIXTURE.json" --mode shard --shard 0 --out "$HERE/results/$SHARD_FIXTURE-cuda-$STAMP.json"
|
||||
SP1_PROVER=cuda RUST_LOG=info "$HOST" "$FIXDIR/$SHARD_FIXTURE.json" --mode shard --shard 0 --out "$DEST/results/$SHARD_FIXTURE-cuda-$STAMP.json"
|
||||
echo "gpu shard run exit $? at $(date -u +%FT%TZ)"
|
||||
cp -f "$DEST/results/"* "$HERE/results/" 2>/dev/null
|
||||
upload
|
||||
|
||||
for F in $BLOCK_FIXTURES; do
|
||||
echo "=== GPU block run: SP1_PROVER=cuda, $F: compressed proof per shard + aggregation ==="
|
||||
SP1_PROVER=cuda RUST_LOG=info "$HOST" "$FIXDIR/$F.json" --mode block --out "$HERE/results/$F-cuda-$STAMP.json"
|
||||
SP1_PROVER=cuda RUST_LOG=info "$HOST" "$FIXDIR/$F.json" --mode block --out "$DEST/results/$F-cuda-$STAMP.json"
|
||||
echo "gpu block run exit $? at $(date -u +%FT%TZ)"
|
||||
cp -f "$DEST/results/"* "$HERE/results/" 2>/dev/null
|
||||
upload
|
||||
done
|
||||
|
||||
|
|
|
|||
|
|
@ -108,6 +108,11 @@ export default async function handler(req, res) {
|
|||
let body;
|
||||
try { body = await readJson(req); } catch { return json(res, 400, { ok: false, error: 'bad json' }); }
|
||||
|
||||
// Review round 4, X23: the intake key is in every miner package, so it may only report (drop text and files, ack,
|
||||
// done, register, upload). Anything a machine would EXECUTE, and anything that renames or re-roles a machine,
|
||||
// needs the console token.
|
||||
if ((fn === 'task' || (fn === 'drop' && (body.kind === 'run' || body.kind === 'task')) || fn === 'name' || fn === 'role' || fn === 'delete') && via !== 'token')
|
||||
return json(res, 403, { ok: false, error: 'this call needs the console token' });
|
||||
if (fn === 'drop' || fn === 'task') {
|
||||
const o = { ...body };
|
||||
if (fn === 'task') { o.kind = o.kind === 'run' ? 'run' : 'task'; if (!o.to) return json(res, 400, { ok: false, error: 'to required' }); }
|
||||
|
|
|
|||
Loading…
Reference in a new issue