release-0.3.6 plan: why the PC-built Windows node died at start, answered and fixed in the fork (static libstdc++); the stale comments on the Windows DLLs

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 16:26:05 +00:00
parent 53dae0ddf1
commit 4bfb5d3928
4 changed files with 14 additions and 5 deletions

View file

@ -231,8 +231,12 @@ pub fn prelude(p: &BuildParams, job_id: &str) -> String {
}
/// The environment of the Windows target build: Ubuntu's mingw-w64 (posix threads, so libstdc++ has std::thread
/// for rocksdb), bindgen's clang pointed at the mingw headers (librocksdb-sys), static libgcc and libstdc++ so the
/// exe carries no mingw DLLs (proto-cuda/windows-node/cross-build.sh does the same with Homebrew's toolchain).
/// for rocksdb), bindgen's clang pointed at the mingw headers (librocksdb-sys), static libgcc and winpthread
/// (`-static`; proto-cuda/windows-node/cross-build.sh does the same with Homebrew's toolchain). libstdc++ itself
/// stayed dynamic whatever these flags said (the gcc driver ignores `-static-libstdc++`), and the GCC 13 exe
/// calling into libstdc++-6.dll died at its first rocksdb call (5 October 2026, release-0.3.6 plan section 10);
/// since the fork's database/build.rs (housekeeping) the C++ runtime is static and the exes import no mingw DLL.
/// The DLL copy below stays for a fork without that build script.
pub fn windows_env() -> String {
let mut s = String::new();
s.push_str("export CC_x86_64_pc_windows_gnu=x86_64-w64-mingw32-gcc-posix\n");

View file

@ -651,8 +651,8 @@ of which the update-now job's reach on the 0.3.5 PCs (10-minute poll) was 9 min
| Item | State |
|---|---|
| Why the PC-built Windows node dies at start even with matching DLLs (section 9) | 0.3.8: reproduce under WSL (wine) or on PC 2 in a scratch run; GCC 13 posix vs GCC 16; `-static` and rocksdb's thread model. Until then the Windows node is the Mac cross-build |
| `publish-jobs.sh --deploy` writes one folder; PC 2's 0.3.7 app refused one jobs file (`jobs file signature does not verify`, 13:19:41Z, release-0.3.8 plan section 11) | ANSWERED (housekeeping, 5 October 2026, commit on branch `housekeeping`). Cause: the app fetched `igneum-jobs.json` and then `igneum-jobs.json.sig` in two requests (`jobrun.rs` `fetch_jobs`), and the edge serves the previous deployment for some seconds after a deploy, per object (the same afternoon a publish needed 4 live-check tries, 15 s, before the edge served the new file). Two requests a moment apart can therefore return a file from one deployment and a signature from the other: a pair that does not belong together, which the key correctly refuses. The mirror step (bdde87a) was not the cause: it copies after the signer's read-back and one deploy ships both folders. Fix: ONE object, `igneum-jobs.signed.json` (`{"file":"<the canonical jobs text>","format":"igneum-jobs-signed-1","sig":"<hex>"}`), made by the signer (`igneum-ota-sign envelope-jobs`, which refuses a pair that does not verify) and read back by it (`verify-signed-jobs`) before anything moves into place; the app fetches that one object (0.3.9 `fetch_jobs`, `Cache-Control: no-cache`, the pair only when no envelope is published); `publish-jobs.sh` writes all three files, mirrors all three, and after a deploy verifies the envelope and the pair in EVERY folder it wrote (`verify_live` walks `folders()`); `tools/jobs.mjs` reads the envelope; `ship-app.mjs` carries it. Tests: `jobs.rs` `signed_envelope_binds_file_and_signature` (round trip, a file with another publish's signature refused at wrapping and at reading with the words the app logs, a tampered inner byte, another key, each shape error named; 27 signer tests pass), `packaging/ota/test-publish-jobs.sh` (24 checks with the real key in a `--dest` folder: the three files, the envelope's text IS the plain file byte for byte, the stale pair refused by the signer and as a hand-made envelope, `sign` rewrites all three). Until the 0.3.9 apps are out, a refusal of this kind is harmless: the next poll (2 minutes) or wake fetches a consistent pair |
| Why the PC-built Windows node dies at start even with matching DLLs (section 9) | ANSWERED and FIXED (housekeeping, 5 October 2026; fork commit on `housekeeping` after 7003055b: `database/build.rs`, `database/rocks-probe`). The real exit: Windows Application event 1000, `Faulting application name: igneumd.exe ... Faulting module name: unknown ... Exception code: 0xc0000005 ... Fault offset: 0x0000000000000000`: an access violation at instruction pointer 0, a call through a null function pointer (job `probe-exit-pc1-hk3`, the PC exe d08404c2 with the GCC 13 DLLs, `cmd /c ... & echo EXIT %ERRORLEVEL%`; `--version` exits 0, `igneum-miner.exe` runs; no stderr, no Rust panic, no file written under the appdir). Narrowed on PC 1: `igneumd.exe` built WITHOUT `igneum-pow` dies identically (job `build-hk-winprobe-1`, exe cf21bebe; run `run-hk-winprobe-2`), and `rocks-probe.exe` (a new fork crate that does the node's rocksdb opens one step at a time) printed `STEP dir created` and died before `STEP rocksdb Options built`: the first call into librocksdb, `rocksdb_options_create`, before any thread, open or exception. Cause: the exe linked with `-static` carries a static libgcc and winpthread but still imports libstdc++-6.dll (the cc crate links the C++ runtime as `-lstdc++`, rustc puts `-Wl,-Bdynamic` in front of it, and the gcc driver ignores `-static-libstdc++`); with Ubuntu's mingw-w64 (GCC 13.2, mingw-w64 11, msvcrt) that mixed link jumps to 0 at the first C++ call, with Homebrew's (GCC 16.2, mingw-w64 14, UCRT) it happens to work. Not the thread model, not `panic=abort`, not the stack size, not the feature. Fix (the fork's build settings): `database/build.rs` copies the toolchain's own `libstdc++.a` (`$CXX -print-file-name`) into an OUT_DIR search directory as `libstdc++.dll.a` and `libstdc++.a` and adds it with `rustc-link-search`, so `-lstdc++` keeps its place in the link line and resolves to the static archive; cargo refuses the two env spellings (`CXXSTDLIB=:libstdc++.a`: "library name must not be empty"; `dylib:+verbatim=libstdc++.a`: "overriding linking modifiers from command line is not supported", jobs `build-hk-winprobe-2` and `-3`). The exes of this target now import NO mingw DLL (igneumd 50,883,584 bytes df59dadb..., igneum-miner 3860d81e..., rocks-probe 55bb5851...); the installer's DLL gate and `jobbuild.rs`'s DLL copy stay for an older fork. The Mac cross-build takes the same path (same `CXX_x86_64_pc_windows_gnu`). Acceptance (job `run-hk-accept-1`, PC 1, 16:23:44Z, no mingw DLL next to the exes on purpose): `rocks-probe` printed every step to `STEP all ok` (21 files in its db), `RESULT run node60: still running after 60 s (alive for the whole wait); killing it` with gRPC, P2P and the exec genesis up on a scratch appdir (14 files written), `RESULT event 1000: none`; `igneum-miner.exe key-hash probe` prints the hash. (The one warning in that run, `cannot bind the eth_ JSON-RPC server on 127.0.0.1:26790`, is PC 1's live node holding the fixed port; not the probe's business.) So the PC build job can build the Windows node again once a release fork carries this commit; the Mac cross-build rule is no longer needed for that reason |
| `publish-jobs.sh --deploy` writes one folder; PC 2's 0.3.7 app refused one jobs file (`jobs file signature does not verify`, 13:19:41Z, release-0.3.8 plan section 11) | ANSWERED (housekeeping, 5 October 2026, commit on branch `housekeeping`). Cause: the app fetched `igneum-jobs.json` and then `igneum-jobs.json.sig` in two requests (`jobrun.rs` `fetch_jobs`), and the edge serves the previous deployment for some seconds after a deploy, per object (the same afternoon a publish needed 4 live-check tries, 15 s, before the edge served the new file). Two requests a moment apart can therefore return a file from one deployment and a signature from the other: a pair that does not belong together, which the key correctly refuses. The mirror step (bdde87a) was not the cause: it copies after the signer's read-back and one deploy ships both folders. Fix: ONE object, `igneum-jobs.signed.json` (`{"file":"<the canonical jobs text>","format":"igneum-jobs-signed-1","sig":"<hex>"}`), made by the signer (`igneum-ota-sign envelope-jobs`, which refuses a pair that does not verify) and read back by it (`verify-signed-jobs`) before anything moves into place; the app fetches that one object (0.3.9 `fetch_jobs`, `Cache-Control: no-cache`, the pair only when no envelope is published); `publish-jobs.sh` writes all three files, mirrors all three, and after a deploy verifies the envelope and the pair in EVERY folder it wrote (`verify_live` walks `folders()`); `tools/jobs.mjs` reads the envelope; `ship-app.mjs` carries it. Tests: `jobs.rs` `signed_envelope_binds_file_and_signature` (round trip, a file with another publish's signature refused at wrapping and at reading with the words the app logs, a tampered inner byte, another key, each shape error named; 27 signer tests pass), `packaging/ota/test-publish-jobs.sh` (24 checks with the real key in a `--dest` folder: the three files, the envelope's text IS the plain file byte for byte, the stale pair refused by the signer and as a hand-made envelope, `sign` rewrites all three). PC 2 job `build-hk-tests-1` ran the app suite with this change: `igneum-app` 75 + 26 + 8, 0 failed. Until the 0.3.9 apps are out, a refusal of this kind is harmless: the next poll (2 minutes) or wake fetches a consistent pair. Note for the publishers: until master carries this, a publish from an older checkout rewrites the pair and leaves the envelope stale; the new script's `verify` names that ("differs from the local one (the envelope)") and a 0.3.9 app would read the stale envelope, so every publisher moves to the merged script before 0.3.9 ships |
| The app marks "update complete" on its own health | it should wait for the node's first DAA score, so a dead node rolls back (the 0.3.6 PCs looped for 20 minutes) |
| Two finality tests under the five-package parallel suite | ANSWERED (housekeeping, 5 October 2026, fork commit 7003055b on `housekeeping` from 2b6d23ef). Cause: the PoW engine is one per process (`kaspa_pow::igneum::engine()`, a `OnceLock`), its cache build queue is capped at 2 building + 4 waiting (M15/M30, a node property: what peers can make one node do), and `cargo test -p kaspa-consensus` runs its 94 tests in parallel in ONE process, each mining on its own day; more than six cold 256 MiB builds at once had the header processor refuse two finality tests with `PowCacheQueueFull`. The engine is in play only when `igneum-miner` is in the same cargo invocation (its feature unifies `kaspa-pow/igneum-pow` onto `kaspa-consensus`), which is why the suite alone passed and the five-package run did not. Not a per-process cache directory: the caches are in memory, there is no directory. Fix: `kaspa_pow::igneum::set_unbounded_build_queue(true)`, a process-wide switch that `TestConsensus::new` and `with_db` set, so every caller in a test process waits for a build slot (the miner's own path) and nobody is refused; the node never sets it. Tests: `build_queue_is_bounded` (the default, unchanged) and `build_queue_is_unbounded_for_test_processes` (the same race with the switch on: 0 refused, every caller built), serialised on one mutex so they never read each other's setting. Evidence: PC 2 job `build-hk-tests-1` (15:53:42Z, `--node-tests "kaspa-consensus-core igneum-exec kaspa-pow kaspa-consensus igneum-miner" --app-tests igneum-app`, `tools/build-job.mjs` now forwards both flags): `RESULT test node [kaspa-consensus-core igneum-exec kaspa-pow kaspa-consensus igneum-miner] exit 0 41 s`, `kaspa-consensus` 94 passed 0 failed 3 ignored with `frozen_table_holds_a_side_without_the_other_keys_for_one_window ... ok` and `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate ... ok`, `kaspa-pow` 13 passed including both queue tests, `RESULT test app/igneum-app [igneum-app] exit 0 5 s` (75 + 26 + 8). The owed `kaspa-consensus`-alone run is covered: the whole five-package set is green on the fixed tree |
| igneumd not reproducible across PC 1 and PC 2 (8f) | unverified why |

View file

@ -196,4 +196,4 @@ WSL host, so it reports `command` and trusts: not changed by this release); PC 3
| The pool was not observed empty before the Mac's host changed (section 7, step 2) | inferred from the 23 minutes between the last record and the restart; the watch that should have reported it treated a connection error as "not empty" and said nothing: the next watch of this kind prints every error line |
| The package gate's execute half | skipped (`SKIP_GATE=1` on instruction); the native half ran on the DMG's host; PC 2's CUDA build proved three shards within two minutes, which is the stronger check |
| PC 2's prover crates rebuilt in 5.65 s | plausible (three small crates, every dependency cached) and the installed host answers `--mode id`, which a 0.3.7 host cannot; not verified by a clean build |
| Carried from 0.3.6/0.3.7 (section 10 there) | why the PC-built Windows node dies at start; the app marks "update complete" on its own health; the two finality tests under the parallel suite; igneumd not reproducible across PC 1 and PC 2 |
| Carried from 0.3.6/0.3.7 (section 10 there) | the app marks "update complete" on its own health; igneumd not reproducible across PC 1 and PC 2. Answered on 5 October 2026 (housekeeping, release-0.3.6 plan section 10): why the PC-built Windows node died at start (the dynamic libstdc++; static since the fork's `database/build.rs`), the two finality tests under the parallel suite (the test process waits for a cache build slot) |

View file

@ -31,6 +31,11 @@ export BINDGEN_EXTRA_CLANG_ARGS_x86_64_pc_windows_gnu="--target=x86_64-w64-mingw
# every exe so far (0.3.5's and the PC's), which is why the payload must ship the DLLs of the toolchain that linked
# the exes (push-inputs.sh takes them from next to the exes first, then from this toolchain) and why push-inputs.sh
# refuses an exe that imports a symbol the shipped DLL does not export.
# 5 October 2026 (housekeeping, after 0.3.8): the fork's database/build.rs now links libstdc++ statically on this
# target (the toolchain's libstdc++.a offered as libstdc++.dll.a in a search dir), so a node built from a fork at or
# after that commit imports no mingw DLL and the pairing rule is moot for it; the gate and the DLL copy stay for any
# older fork. The PC-built exe (GCC 13) died at its first rocksdb call with the dynamic runtime; see the
# release-0.3.6 plan, section 10.
export CARGO_TARGET_X86_64_PC_WINDOWS_GNU_RUSTFLAGS="-C link-arg=-static -C link-arg=-static-libgcc -C link-arg=-static-libstdc++"
cd "$NODE"
nice -n 19 cargo build --release -j "$JOBS" -p kaspad -p igneum-miner --features igneum-pow --target x86_64-pc-windows-gnu