igneum/docs/plans/shard-test-pc2.md

6.8 KiB

Shards ready to test on PC 2: the signed job channel

4 October 2026. the project lead's ask: shards ready to test on PC 2 (RTX 5090, WSL2 Ubuntu 24.04, the prover toolchain installed this morning, the Igneum Miner app running as machine 1ccfe586) while he is away, and from then on "one app on both PCs that you can send over the line commands to and files so we can continually test everything and build without input". PC 2 has no Claude session and nobody at the keyboard, so the line is the app itself: a signed job channel next to the over-the-air manifest. The relay (relay/) stays for the Mac and for humans; its PC agent is replaced.

What is built (this commit)

Piece Where State
Job model: parse, Ed25519 verify (OTA key), targeting (machine id 8 or 16 hex or all, platform, requirements), expiry, once-per-id ledger, helpers app/igneum-app/src/jobs.rs 7 unit tests pass on the Mac
Runner: 10-minute poll, requirement probes (wsl, wsl-prover, nvidia), the kinds run, fetch, collect, restart, update-now, shard-benchmark, reports to the log intake as job-<id>-<machine id8> with a SUMMARY {json} first line, 5-minute progress reports, time caps, abort on quit app/igneum-app/src/jobrun.rs compiles for macOS and x86_64-pc-windows-gnu; not yet run on a PC
Engine hooks: Cmd::Job, miners stopped and held for a job, node restart, app relaunch, updater on demand; no OTA apply under a running job app/igneum-app/src/engine.rs
Dashboard: the job strip (running: "Job: shard benchmark running, N min" plus the RESULT lines as they arrive; afterwards the outcome for 30 minutes), Settings: "Allow remote jobs from Igneum (signed)" ON by default with the key fingerprint, Check now, the history table (id, kind, started, exit, report) app/igneum-app/ui/
Signer: sign-jobs, verify-jobs (refuses a file with a bad job) app/igneum-app/src/bin/ota-sign.rs run with the real key: signs, verifies, refuses a tampered file and a bad job
Publisher: `publish-jobs.sh add list remove
Reader: the published file (signature checked), status per machine, a job's result, watch tools/jobs.mjs reads the live intake (no job reports yet)
Relay playbook, the same run in its relay form (reference, and the model for a run job) relay/playbooks/shard-test.ps1 parse-checked by windows.yml (folder added to the check)

The prove package: ~/Desktop/igneum-prove-wsl2.zip (286,438 bytes, sha256 5e7b56f5...950373) is byte-identical to the hosted dl/<token>/igneum-prove-wsl2.zip, and its contents match proving/windows-wsl2/, proving/igneum-prove/ (sources and Cargo.lock) and proving/fixtures/ (338, 341, 344 and the three test fixtures); the only difference is the thiserror pin make-package.sh writes into evm-types. No rebuild was needed.

The first job (after 0.3.2 is on PC 2)

The app on PC 2 is 0.3.1 today. The job channel is in the build after it: the project lead cuts 0.3.2 with the miner fix, the OTA installs it, then:

packaging/ota/publish-jobs.sh add --kind shard-benchmark --target 1ccfe586 --title "Shard proof run on the 5090" --deploy

This hosts the zip (already there), signs the file, deploys the downloads folder and verifies the live file. Within 10 minutes (or at once from Settings > remote jobs > Check now) PC 2's app: stops its miners (the node keeps running), waits for the GPU under 5%, downloads the zip into %LOCALAPPDATA%\igneum\prove, extracts it fresh, runs wsl -d Ubuntu-24.04 -- bash /mnt/c/.../prove-shard.sh block-338-shard1 "block-341-shards2 block-344-shards4" under a 90-minute cap, uploads every RESULT and STAGE line, the results/*.json and the prove log, and restarts the miners. Options: --fixtures "...", --cap-minutes, --wsl-user [user] (when [user] is not the distro's default user), --distro.

Watching from the Mac:

node tools/jobs.mjs                      the published file, signature checked
node tools/jobs.mjs status               the latest job per machine with its SUMMARY line
node tools/jobs.mjs watch <job id>       every 30 s until final; the result files land as labels result-<name>
node tools/logs.mjs                      the app's own uploads (the engine log shows the job's events too)

The RESULT lines fill the empty GPU column of the bench-log skeleton ("shard proving on the RTX 5090"). The first S_p point comes from the shard stage times (docs/plans/proving-v0.md).

Everything else over the same line

publish-jobs.sh add --kind run --target 1ccfe586 --script x.ps1 [--elevated] [--stop-miners] [--timeout-minutes 60]
publish-jobs.sh add --kind fetch --target all --file ~/Desktop/thing.zip --dir prove --extract --fresh --extract-dir thing
publish-jobs.sh add --kind collect --target all --glob "logs/app-*.log" --glob "prove/igneum-prove-wsl2/results/*.json" --command "nvidia-smi"
publish-jobs.sh add --kind restart --target ae432dc7 --what miners|node|app
publish-jobs.sh add --kind update-now --target all

A machine runs an id once; the same thing again is a new add. Jobs expire (48 hours by default) and are dropped from the file on the next write. Everything a job writes stays under the app data folder except what a run script does, which is the operator's responsibility. elevated on an unattended PC fails after the UAC prompt times out: that needs someone to click, or UAC set to elevate without prompting.

Not verified from the Mac

What Why How it gets checked
The runner on Windows: wsl.exe output through the sink, taskkill /T on the cap, the elevated wrapper no PC reachable from this session the first shard-benchmark and a run job on PC 2 after 0.3.2; the engine log (app-*.log) carries every step
That [user] is the distro's default user (the probe and prove-shard.sh run as the default user unless wsl_user is set) not visible from here if the probe fails with cargo missing, re-add the job with --wsl-user [user]
The dashboard strip and the Settings block in a real window the dashboard needs the engine's API; only the JS parse and the state shape were checked the first job on either PC
The 90-minute cap against a cold build (first run compiles 10 to 30 minutes, approximate) the toolchain was pre-built this morning in ~/igneum-prove, and prove-shard.sh rsyncs the same sources over it, so the build should be incremental the STAGE timestamps in the report
The relay: PC 2 is not registered (both PCs report the hostname DESKTOP-KMCV30N; the relay's PC1 is whichever started igneum-agent.bat), so no relay task was queued by the project lead's change of plan the relay agent is not the channel for the PCs the playbook stays as the relay form