diff --git a/docs/fud-ledger.md b/docs/fud-ledger.md index da25983a..83202996 100644 --- a/docs/fud-ledger.md +++ b/docs/fud-ledger.md @@ -296,6 +296,8 @@ Evidence: design doc Finality v2, Residual risks bullet 3. Fix: overclaims list, Sweep (5 October 2026, evening): stated. `site/litepaper.html`, Finality ends with the pool sentence (overclaim 33); Governance opens with "governed by the hashrate that powers it", pool concentration named as the governance risk with the devnet measurement, top-3 keys 34.5% of 8,090 blocks (ledger sweep, X14 run). Overclaim 43. +Narrowed (7 October 2026, the pool lane, mission item 11): a conforming pool holds no vote. Pool-0 (`pool/`, spec 09) names the member's own vote key in every header it issues and holds no key of its own, so its hashers' weight is theirs (measured on the private fast-time runs of 5 and 7 October: every pool block carried the finding member's key); and the open pool (spec 09 section 9.12, `igneum-pool --open`) has no operator at all: every member builds its own templates from its own node, and a block pays the window from its own coinbase by the executor's split rule. The litepaper's sentence stays true of a custodial pool (`vote_mode: pool`, refused by the member's client by default and visible on chain as one key per payout address); the next litepaper pass should say "a custodial pool" where it says "pools". + ### F11. VDFs are exotic "A class-group VDF with Wesolowski proofs in a consensus-critical path, in a project with no cryptographer. Chia needed years and still got timelord ASICs." @@ -655,6 +657,8 @@ Answer: Correct. Stratum v2 job declaration lets a hasher choose transactions wh Evidence: none in repository. Fix: overclaims list, item 45. +Narrowed (7 October 2026, the pool lane): on Igneum the vote key in a pool's header is the member's, not the pool's (spec 09 section 9.6, built in `pool/` since 5 October and in the open pool since 7 October), so the vote-key caveat no longer applies to a conforming pool. Transaction choice: still the pool's in mode A (mode C, the member's own template, is designed and not built); in the open pool the member's own node chooses the transactions, since the member builds every template, which is Stratum v2's job declaration without an operator to decline it. The litepaper's "pools can be bypassed on transaction choice" stays, with "and in the open pool there is no pool to bypass" when it is next edited. + Cross-reference (external review, 3 October 2026, night): whether members use declared templates is measured under O-9.5; concentration reporting is X14. Sweep (5 October 2026, evening): stated. `site/litepaper.html`, Governance bullet "Pools can be bypassed on transaction choice", a pool may decline, vote keys stay with the pool, label Designed with spec section 9 (overclaim 44). diff --git a/docs/plans/pool.md b/docs/plans/pool.md index 30c21215..e39e3843 100644 --- a/docs/plans/pool.md +++ b/docs/plans/pool.md @@ -236,3 +236,262 @@ a wrong-size cache and mismatched silently); a pool on a class v4 network needs Building the pool crate on igneum-build-1: the pool reads the fork through the `vendor/igneum-node` symlink; lib.sh syncs the fork worktree it points at (`vendor/igneum-node-pr`) as a whole repository, and the symlink itself is created once on the box by hand (`ln -s igneum-node-pr /srv/builds//vendor/igneum-node`), the one step the scripts do not do. + +### 9.1 7 October 2026: the daemon refuses to start half-alive + +The fleet agent's night (6 October, 23:01Z to 23:32Z): the rebased pair never ran as a test. The daemon was started while +the node still held its member port, then four members (the wind-down's weight rule freed four pods, not ten) sat +connected with no job while every template fetch timed out; nothing in the daemon's own log said so except one line per +member. The row from it, now code (pool `269364a`+): the member and API listeners are bound in `main` before anything +else and a port that cannot be bound ends the start with exit code 2 and `POOL NOT STARTED: cannot bind the members +listener on : `; the node must answer a template with its `pow_epoch` before the pool serves anyone, +retried for `--node-wait-secs` (default 60, each attempt logged), else exit code 3; a `STATUS` line every 30 s +(`members= jobs_issued= shares_accepted= shares_rejected= blocks= template_failures=; node ok | NODE NOT ANSWERING | +NO TEMPLATE YET`). Test: `server::bind_tests` (a held port refused with the address in the message, the freed port binds). +The 10-member run is still owed: a clean 30-minute window, the daemon bound and its `node ... answers templates` line +printed before the first member, on the next pods the weight rule frees (the table must read over 75 percent signed first). + +### 9.2 7 October 2026, 06:00Z to 06:33Z: the re-run (five members), what it settled and what it found + +Settled: the class question. pm-1 and pm-2 (RTX 3070, grpc url `none`) at +7 min: 36,917 and 34,323 shares accepted, 0 +rejected, 0 WORKER MISMATCH, 0 refused, on the class v3 program of the pool's node's epoch (`5530a50d...`, era = the devnet +genesis hash), the genesis day (20729) and dataset size (2^28) installed from the daemon's `seeds` line. The 6 October +failure does not reproduce. + +Found: the daemon stalls on a real chain. From 06:01Z (the first block found) no job was issued for seven minutes, pm-3 +and pm-4 never got one, every block found on the stale templates (55, then 105, 480 DAA behind the virtual) was an +orphan, vardiff could not act (a new share target rides with a job) so the two hashing members sent about 90 shares a +second each and the share check read 9.8 ms on pool-1's cores (1.35 ms on the Mac); after a swap to a 20 s template +timeout (f808b3f3) with the data dir kept, the stall returned within a minute. The node built templates in 0.6 ms all +the while (its prewarm line) and served its own solo miner. + +The cause, from the daemon's code: `confirm_loop` restarted its chain walk from the PRUNING POINT on any failed +`getVirtualChainFromBlock`, and then called `get_block` for every chain block since (about 120,000 on the devnet) over +the ONE gRPC connection the templates used (the client is one request stream per connection, 5 s request timeout); one +timed-out request under four parallel template fetches started the walk, every later request waited behind it and +timed out, which restarted the walk again. The kept state file (105 pending blocks) restarted it after the swap. The +private measurement run (section 5) never saw it: a 600-block chain walks in a moment. + +Fixed (pool commit after bd49c2a9): the walk and the network numbers run on a second gRPC connection (`pool.walker`); a +failed chain call keeps its cursor (moved to the sink only when the node no longer knows it, never to the pruning +point; test `node::walk_tests`); a tick walks at most 600 chain blocks and says how many; a start with pending blocks +walks from the sink and says that older ones resolve by the orphan rule. `--template-timeout-s` stays (default 20). + +Still open from the night, for the next window: the node log showed a new gRPC connection about every 0.8 s with the +count steady at 5 (something opens and closes one each time; the daemon's re-subscribe loop is the suspect, the +subscribe-line count in the daemon log decides it); vardiff's new target should apply to the member's current work +without waiting for a job (a one-line member change, needs a miner rebuild); the 9.8 ms share check on the pool box's +cores against 1.35 ms on the Mac (the verifier's day cache is 1 GiB and random reads on a rented box's memory are the +cost; two verify threads saturate at about 200 shares a second, which vardiff must keep far away); and the 10-member +load figure itself, not taken. + +Consequences per tier: a pool operator needs a box whose memory serves the 1 GiB cache at speed (a rented 2 vCPU box +verifies about 100 shares a second per core at 9.8 ms; at one share per 10 s per member that is 1,000 members per +core, so the verifier is not the limit once vardiff holds); a member on any card is unaffected by any of this; the +chain walk now costs the node at most 600 `get_block` calls per 5 s on its own connection. + +### 9.3 The re-run's close (06:33:33Z) and the member rows from its logs + +After 2f6c0358 (06:20:34Z to 06:33:33Z, 13 minutes, four members): 207 blocks found, 144 confirmed, 69 orphaned (13 +inherited pending at the swap, 56 on its own jobs: a 27 percent residual at 1 bps), 20,404 shares accepted, 0 rejected, +0 mismatches, 0 refused, every STATUS `node ok`, the gRPC connection churn gone (ids #3 to #5 once each in a minute; the +193 connections in 150 s were under f808b3f3's pruning-point walk), re-subscribe lines 2 (one per daemon start: the +re-subscribe loop was never the churn). Before it, f808b3f3 had worked through the walk by about 06:14Z: 245 blocks, +69 confirmed, 163 orphaned. First CONFIRMED line 06:24:32Z, 4.80 IGN to the finder. Logs: `~/Desktop/fleet/pool-run-0607/`. + +The residual orphans are template latency: the daemon's STATUS reads `last template 1 s ago` and the solo miner on the +same node reads `template_ms=1632`; pool-1's node answers a template request in 1 to 1.6 s although its mining manager +answers a per-member request from its cache by a coinbase rewrite (`modify_block_template`) and its prewarm builds in +0.6 ms. At 1 bps a template that arrives 1.6 s old loses about a quarter of the blocks found on it, pool or solo. That +latency is the node's RPC path, a row for the node lane (measure `get_block_template` round trips on a devnet node under +a miner and a pool; the suspect is the RPC service's `pow_epoch` derivation per request). 82 "RPC request timeout" lines +after the 2f6c0358 swap are the gRPC client's own 5 s request timeout on that path; `--template-timeout-s` cannot +lengthen it. + +From pm-1's log (1.09 GB): 4,342,515 lines are `worker: error N epoch seed mismatch: this worker holds epoch eea66ce7...`, +one per re-queued job, for the 40 s the worker compiled the new epoch's pack (NVRTC 40,327 ms) and again after every +reconnect, because a session respawned the worker (a second process beside the first, the same compile again) and +the member fed jobs for a pair the worker did not hold yet. Vardiff from the same log: shift 10 to 11 (idle easing +while the worker compiled, which consumed the sized first correction), then 11 down to 0 one step per 30 s over 5.5 +minutes at about 190 shares a second, the load that put the share check at 10 to 11 ms on pool-1's cores. + +Fixed in the member (fork commit after b8070476): one worker process per run, jobs held for a pair whose prepare is +pending and never re-prepared once held (`WorkerMemory`, test `jobs_are_held_while_a_pair_is_being_prepared`), worker +error lines summarised (one per class, then a count per 1,000), `set_target` applied to the current work at once. +Fixed in the pool (`vardiff.rs`): the idle easing no longer consumes the sized first correction (test +`an_idle_easing_does_not_consume_the_sized_first_correction`). + +Consequences per tier: a member on a 3070-class card loses 40 s of hash at every epoch roll to the NVRTC compile unless +its pack is prepared ahead (the pool's `seeds` line carries `next_epoch_seed` for that, the member's prepare-ahead is +the next member row); a member on a reconnect now keeps its worker and its compiled pair; a pool operator's verifier +sees the sized correction within 10 s of a member's first shares (one share per 10 s per member from then on) instead of +5 minutes of a 190-share-a-second flood; the 10-member load figure is still owed and now has a clean form to run in. + +## 10. 7 October 2026: pool-0 finished, TLS, and the open pool (mission item 11) + +Branch `pool-finish` on `release-0.3.19` 44eee05b with the four daemon commits of `pool-v0-rebase` cherry-picked +(the node fork pinned for the pool work is 0.3.19, whose kaspa-pow needs the latency ladder's igneum-pow; master +b92a5fd4 does not carry it, so the branch sits on the release tree); fork branch `pool-finish-node` on +`release-0.3.19-node` dc141409 with `pool-v0-rebase` merged (the pool-mode miner). Ordered by the project lead on 7 October 2026, +10:1x UK, built in the order of mission 2.11: pool-0, TLS and the page rows, the share sidechain. + +### 10.1 Pool-0 + +What 2.11 asks of pool-0 was in v0 (section 2): the member's vote key in every header, the pool with no key and no +vote, PPLNS, a payout round every 60 s, a 1 percent fee. What was missing was the fee published beside the dev fee +(reinvent 3.5): `welcome.share_scheme` now carries `software_dev_fee_percent` (0 in pool mode: the pool's fee is the only +fee), the page's Fees row says both, and `site/miner.html`'s dev-fee card says pool-0 charges the same 1 percent so solo +and pool cost the same and the choice is about variance alone. The jobs and seeds lines carry the latency ladder's rung +(`shadow_reps`) beside the class and the era, which 0.3.19's program needs; a member checks it against its own node as it +checks the seed, the class and the era (a pool that could choose the rung could choose the program). + +### 10.2 TLS (spec 9.3, O-9.7 closed, Q67) + +`pool/src/tls.rs`: rustls 0.23 with ring, TLS 1.3 only. `--tls-cert`/`--tls-key` (a chain from a trusted root) or +`--tls-self-signed` (a P-256 pair made under the data dir on first start, read back afterwards; the pin +`BLAKE2b("igneum-pool-cert-pin-v1" || DER)` printed at start, shown on the page and in `/api/stats`). The member +(`igneum/miner/src/pool.rs`, fork): `--pool-tls` (the Mozilla roots, the pool's host as the server name) or `--pool-pin +` (a verifier that accepts exactly the pinned certificate). The binding: both sides export 32 bytes with the label +`EXPORTER-igneum-pool-binding` and the chain id (8 bytes LE) as the context; the member signs `igneum-pool-binding-v1/ +|| chain id LE || exporter` under its vote key with its own tag (`DST_BINDING`, `consensus/core/src/finality.rs`); the +pool verifies it against the member's key and refuses the `authorize` with `bye` otherwise. Tests: `finality.rs` +(the signature holds for one exporter and one chain id, is no proof of possession); `tls.rs` (a self-signed pair read +back with one pin, a pinned handshake, both sides' exporters equal, the binding verified on its own connection and +refused on a second, a wrong pin never handshakes). Testnet and mainnet refuse to start in the clear without +`--allow-plain`; the devnet's local daemons stay plain. HiveOS: `pools://` or `POOL_PIN=` (`packaging/hive`). + +### 10.3 The page rows (polish Q68 to Q73) + +| Row | Done | +|---|---| +| Q68 | the page's node line: "node ok, ", "node syncing", "no template", "node unreachable since