diff --git a/docs/design/class-v5-stored-state.md b/docs/design/class-v5-stored-state.md index 19cfb9e7d..fd9b4e3c5 100644 --- a/docs/design/class-v5-stored-state.md +++ b/docs/design/class-v5-stored-state.md @@ -21,6 +21,8 @@ | 09:48 | The known-failed case FAILED as it must: no flip at 7,519 bps of byte 5 (two signalling nodes of three plus the stateless node), 665 blocks, 0 rejected, four sinks equal; the harness's own two faults fixed (e0c555cb) | | 10:00 | The first flip run read a stranger's chain (two nodes of the known-failed case were still bound to its ports); the harness now kills its own previous pids and fails on a bind error, with a port base per case (6c983cba); the miner's engine panicked on a state refusal while the executor had not yet captured the seed block: it now says the refusal once and asks again every two seconds (bdc3bb34); the flip re-run started 10:24 | | 10:23 | The fleet's 4090 pod (cv5-2) measured (section 7): rate equal within 0.01 percent, build +0.24 ms, watts a warm-up drift with the classes interleaved, every vector and sample bit-exact | +| 10:38 | The flip re-run (`--signal 5,5,5 --expect flip --stale 2`, ports 29830 and up): the chain flipped v4 to v5 at epoch 8 by signal on 3 of 3 nodes (`Program class v5 by miner signal: epoch 8 ... weakest 10000 bps`), the first epoch with seven full windows after v4; and node 0's templates timed out from the flip on: the engine's provider took the executor's state lock from inside the template path while the follower held it and waited on consensus (a lock-order inversion), and the miners' stream fetches timed out the same way. Fix: the follower publishes each finalised epoch's stream (its cut passed, the stream rebuilt to its root) into a separate cache the provider and the RPC read without the state lock; the harness points the stateless node's miner at its own node's exec port. The 60x profile's day boundary was crossed inside the earlier runs (days 1244001 to 1244003 in the summaries) | +| 10:4x | The Counter ASIC coordinator relays the project lead's ruling of option A on AP-F8-1: the injecting-source load draw lands in class v4 itself on the hash lane's branch ca3-v4-amend (0.3.19); class v5 takes it by merging that branch once its commit is up, not by a second implementation; the page's section 11 row stands as the record of the bound and its gate | ## 1. The claim, in one paragraph @@ -225,7 +227,7 @@ Per tier, what the 4090 rows mean: every NVIDIA card from 8 to 32 GB keeps its h Implemented on `class-v5` and `class-v5-node` (this page's time line in section 0 names each commit): the design; `igneum_pow::blake2b`, `igneum_pow::state` (leaves, sample, the leaf XOR in `memhard::derive_items`), `V5_CLASS`, `ProgramClass::V5`, generator 5, the emitters' build kernel with the leaf buffer, the CLI's `--program-class v5 --state `; the fork's `DayStream`, the day-state capture in the executor, the provider in `kaspa_pow`, `program_class_v5_activation_daa`, the v5 signal rule and tally, the RPC fields, the daemon lines, the miner's fetch; the tests with the known-failed cases first; the pinned v5 packs; the harness. -Owed, named here so nobody looks for them: an acceptance-rule bound on a program's hot-set share (the Counter ASIC lane's item for v5, 09:5x UK, on main's word: the attack-pass lane's F8 found that on a class v4 program the top 0.1 percent of items take 0.52 percent of reads, 4.05x a uniform map, one item 153x the mean, most of it class v4's designed per-site windows of layer 8; the v5 rule rejects a draw whose windows and offsets coincide into a hot set beyond the window model's tail, with the rejection's cost stated in accepted programs per 64 seeds under the 2.0 bounds; the bound's definition and number come from the hash lane's window-model analysis on branch ca3-v4-uniform and the gate is F8's 64-seed census, `tools/attack/f8-uniform` on branch attack-pass phase E, the top 0.1 percent within the bound on every seed, re-gated by the attack-pass lane; nothing of it touches v4, which is on the vote; the bound landed 10:3x UK from `docs/analysis/ca3-v4-uniform.md` 095f84a7: the fault is a lossy-sourced load, a load whose source register was last written by an entropy-losing op (or, mul, mulhi), whose saturated value recurs at (3/4)^32 per read and lands on one item through the era map; 96.6 percent of today's class v4 programs carry one; the bound H = W_0.1 from the 16 window draws (0.115 to 0.251 percent) plus the sum over load sites of h(last writer), h = 0.30 percent for or, 4.5 for an or chain, 0.067 mul, 0.049 mulhi, 0 for an injecting op or a rotate, with H at or under 1.2 x W_0.1, equal to the static rule "every load's source was last written by an injecting op or a rotate"; as a rejection it would cost 96.6 percent of candidates, so the form for v5 is a generator draw: a load's source drawn from the registers whose last writer injects, no attempts lost, rule (a) unchanged; its worth to a chip today, at most 1.005x on about half the hours, 1.048x on 5 percent, 1.067x at the ceiling; it changes the v5 program draw at the load sources alone, so "v4's draw for draw" in section 2 becomes "v4's draw but for the load sources" once it lands, with the pinned v5 pack re-exported; gate as above); beside it, the mixer-draw acceptance rule of the attack-pass lane's AP-F4-1 (`docs/analysis/attack-pass/f4-weakday.md` sections 6.7, 7 and 8 on branch attack-pass, relayed 10:0x UK): the day key's MUL block has a tail of cheap multipliers (low NAF weight), a tail of a sum with no weak class behind it; against the DSP-bound datapath F4 passes (0 of 2^28 days over 1.1x), on the LUT-adder metric 5,476 of 2^24 days are over 1.1x (the worst 1.173x; the worst calendar day in the first 100 years is chain day 29,337 at 1.121x), worth at most 12.1 percent more hash rate that day to a per-day LUT-recompute FPGA and nothing to a stored-dataset FPGA or any chip; the v5 rule for the day's mixer draw: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values, NAF weight at least 4 per word, at least 4 distinct ROT amounts, total rejection 6.1e-4 per day, the first calendar redraw at day 22,633 (5.2 years in), no devnet or testnet pack changes; gate: F4's harness (the 2^24-day census under both metrics, the worst day's M1 cost at or above 211 after the rule), re-gated by the attack-pass lane against this branch; it is the mixer-draw side of `memhard.rs` (`MixParams::with_shape`), independent of the leaves, and lands behind the v5 class; the exec snapshot wire carrying the day streams (version 2); the GPU worker hosts' leaf buffer (`proto-cuda/nvrtc`, `proto-opencl`, `proto-metal`: the kernel text carries it, the hosts must upload it); the pool protocol's daily fetch (spec 09); the day-state witness in the pruning-proof format (with the class-signal witness); the spec text (01 1.8.5 the leaf line, 1.12 the cut and the reference block, 10 the witness); the 2019-class core row (O-1.14); Devnet 2 across a real day boundary. +Owed, named here so nobody looks for them: an acceptance-rule bound on a program's hot-set share (the Counter ASIC lane's item for v5, 09:5x UK, on main's word: the attack-pass lane's F8 found that on a class v4 program the top 0.1 percent of items take 0.52 percent of reads, 4.05x a uniform map, one item 153x the mean, most of it class v4's designed per-site windows of layer 8; the v5 rule rejects a draw whose windows and offsets coincide into a hot set beyond the window model's tail, with the rejection's cost stated in accepted programs per 64 seeds under the 2.0 bounds; the bound's definition and number come from the hash lane's window-model analysis on branch ca3-v4-uniform and the gate is F8's 64-seed census, `tools/attack/f8-uniform` on branch attack-pass phase E, the top 0.1 percent within the bound on every seed, re-gated by the attack-pass lane; nothing of it touches v4, which is on the vote; the bound landed 10:3x UK from `docs/analysis/ca3-v4-uniform.md` 095f84a7: the fault is a lossy-sourced load, a load whose source register was last written by an entropy-losing op (or, mul, mulhi), whose saturated value recurs at (3/4)^32 per read and lands on one item through the era map; 96.6 percent of today's class v4 programs carry one; the bound H = W_0.1 from the 16 window draws (0.115 to 0.251 percent) plus the sum over load sites of h(last writer), h = 0.30 percent for or, 4.5 for an or chain, 0.067 mul, 0.049 mulhi, 0 for an injecting op or a rotate, with H at or under 1.2 x W_0.1, equal to the static rule "every load's source was last written by an injecting op or a rotate"; as a rejection it would cost 96.6 percent of candidates, so the form for v5 is a generator draw: a load's source drawn from the registers whose last writer injects, no attempts lost, rule (a) unchanged; its worth to a chip today, at most 1.005x on about half the hours, 1.048x on 5 percent, 1.067x at the ceiling; the project lead ruled option A at 15:2x UK by the coordinator's relay: the draw lands in class v4 itself (branch ca3-v4-amend, the hash lane, 0.3.19, a new program stream and vectors, the amended class with its own generator stamp), and class v5 inherits it by merging that branch into class-v5 once its commit is up, the pinned v5 pack re-exported on that base; this row is the record of the bound and its gate, satisfied by the v4 amendment; gate as above); beside it, the mixer-draw acceptance rule of the attack-pass lane's AP-F4-1 (`docs/analysis/attack-pass/f4-weakday.md` sections 6.7, 7 and 8 on branch attack-pass, relayed 10:0x UK): the day key's MUL block has a tail of cheap multipliers (low NAF weight), a tail of a sum with no weak class behind it; against the DSP-bound datapath F4 passes (0 of 2^28 days over 1.1x), on the LUT-adder metric 5,476 of 2^24 days are over 1.1x (the worst 1.173x; the worst calendar day in the first 100 years is chain day 29,337 at 1.121x), worth at most 12.1 percent more hash rate that day to a per-day LUT-recompute FPGA and nothing to a stored-dataset FPGA or any chip; the v5 rule for the day's mixer draw: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values, NAF weight at least 4 per word, at least 4 distinct ROT amounts, total rejection 6.1e-4 per day, the first calendar redraw at day 22,633 (5.2 years in), no devnet or testnet pack changes; gate: F4's harness (the 2^24-day census under both metrics, the worst day's M1 cost at or above 211 after the rule), re-gated by the attack-pass lane against this branch; it is the mixer-draw side of `memhard.rs` (`MixParams::with_shape`), independent of the leaves, and lands behind the v5 class; the exec snapshot wire carrying the day streams (version 2); the GPU worker hosts' leaf buffer (`proto-cuda/nvrtc`, `proto-opencl`, `proto-metal`: the kernel text carries it, the hosts must upload it); the pool protocol's daily fetch (spec 09); the day-state witness in the pruning-proof format (with the class-signal witness); the spec text (01 1.8.5 the leaf line, 1.12 the cut and the reference block, 10 the witness); the 2019-class core row (O-1.14); Devnet 2 across a real day boundary. ## 12. The litepaper paragraph and the ledger row diff --git a/infra/fast-time/class-v5-signal.mjs b/infra/fast-time/class-v5-signal.mjs index fd65d2193..2a33464d9 100644 --- a/infra/fast-time/class-v5-signal.mjs +++ b/infra/fast-time/class-v5-signal.mjs @@ -155,7 +155,8 @@ const SIGNALLING = [n0, n1, n2]; for (const n of nodes) log(`n${n.i}: ${n.grepLog(WINDOW_LINE).map(l => l.replace(/^.*?(Program class v4 signal window)/, '$1'))[0] || '(no window line)'} | ${n.grepLog(OWN_LINE).map(l => l.replace(/^.*?(this node signals)/, '$1'))[0] || '(no signal line)'}`); log(`n0 digest: ${n0.grepLog(/Consensus params digest/).map(l => l.replace(/^.*?digest: /, '').slice(0, 16)).join(' ')}`); nodes.forEach((n, i) => { - const extra = n.stateless ? [] : ['--exec-rpc', n.execRpc]; + // the stateless node's miner is pointed at its own node's (absent) exec port, so its refusal is its node's and not a stranger's + const extra = ['--exec-rpc', n.execRpc]; if (STALE === i) extra.push('--freeze-state'); miner(CPU_MINER, ['mine', n.grpc, '1', String(SECS), `cpu${i}`, '--engine', 'igneum-pow', '--payout-label', `cpu${i}`, '--status-secs', '30', '--no-vote', ...extra], `cpu${i}`, { IGNEUM_POW_DAY_MS: String(DAY_MS) }); });