diff --git a/docs/design/class-v5-stored-state.md b/docs/design/class-v5-stored-state.md index cccbae22a..3bc4098ba 100644 --- a/docs/design/class-v5-stored-state.md +++ b/docs/design/class-v5-stored-state.md @@ -27,6 +27,7 @@ | 11:47 to 11:59 | The flip run on the quiet box: the chain did everything the design says (the flip, four v5 epochs, a root per epoch, the stale miner off on 52 of 52, the stateless node stopped at 480); six `FAILED CHECK`s were the harness's own scope (it compared the stateless node's sink with the others, counted the stale miner's rejections as a fault, and read the flip epoch's stream after the executor had dropped it: captures are kept two epochs); the checks fixed | | 12:00 to 12:11 | PASS on every check (section 10) | | 12:12 to 12:23 | The no-flip case PASS (section 10); both branches pushed; master merged for the build-2 test route (d1285cef) | +| 13:2x | The coordinator's correction: the hot-set bound is NOT satisfied by sub-version 1 (F8's re-gate: nine of the first 30 seeds over 1.2x, p31 at 29.3x, 0.19 percent hot-set programs; saturated values pass saturation-preserving writers); v5 inherits the gap by merge and takes sub-version 2 by merge when it lands; the AP-F4-1 and AP-F1-1 rules come in the next round; the lane stops here | | 12:5x to 13:21 | The amended class v4 taken in: ca3-v4-amend 8c728ca3 merged into class-v5 (4e737543) and release-0.3.20-node into class-v5-node (699db5a2); the source rule's class key sets the state flag aside so v5 and its rungs draw under it (a unit test pins the v5 chain draw equal to the amended v4's with no lossy-sourced load, the unamended v3 stream the known-failed case); generator 5 keeps `program_id(5, seed, attempt)` with no sub-version (the hash lane's reading: the suffix separates two streams inside generator 4); class v5 is object byte 6 above the amended v4's 5 (the fork, the harness, the default byte); igneum-pow 67 unit and 39 integration tests on igneum-build-2, the packs test with the re-exported control pack, the flip case PASS again with bytes 6,6,6 (the flip at epoch 8 on 3 of 3, 183 v5 blocks, the stale miner off on 59 of 59, the stateless node stopped at 480, equal roots) | | 10:4x | The Counter ASIC coordinator relays the project lead's ruling of option A on AP-F8-1: the injecting-source load draw lands in class v4 itself on the hash lane's branch ca3-v4-amend (0.3.19); class v5 takes it by merging that branch once its commit is up, not by a second implementation; the page's section 11 row stands as the record of the bound and its gate | @@ -239,7 +240,7 @@ Per tier, what the 4090 rows mean: every NVIDIA card from 8 to 32 GB keeps its h Implemented on `class-v5` and `class-v5-node` (this page's time line in section 0 names each commit): the design; `igneum_pow::blake2b`, `igneum_pow::state` (leaves, sample, the leaf XOR in `memhard::derive_items`), `V5_CLASS`, `ProgramClass::V5`, generator 5, the emitters' build kernel with the leaf buffer, the CLI's `--program-class v5 --state `; the fork's `DayStream`, the day-state capture in the executor, the provider in `kaspa_pow`, `program_class_v5_activation_daa`, the v5 signal rule and tally, the RPC fields, the daemon lines, the miner's fetch; the tests with the known-failed cases first; the pinned v5 packs; the harness. -Owed, named here so nobody looks for them: an acceptance-rule bound on a program's hot-set share (the Counter ASIC lane's item for v5, 09:5x UK, on main's word: the attack-pass lane's F8 found that on a class v4 program the top 0.1 percent of items take 0.52 percent of reads, 4.05x a uniform map, one item 153x the mean, most of it class v4's designed per-site windows of layer 8; the v5 rule rejects a draw whose windows and offsets coincide into a hot set beyond the window model's tail, with the rejection's cost stated in accepted programs per 64 seeds under the 2.0 bounds; the bound's definition and number come from the hash lane's window-model analysis on branch ca3-v4-uniform and the gate is F8's 64-seed census, `tools/attack/f8-uniform` on branch attack-pass phase E, the top 0.1 percent within the bound on every seed, re-gated by the attack-pass lane; nothing of it touches v4, which is on the vote; the bound landed 10:3x UK from `docs/analysis/ca3-v4-uniform.md` 095f84a7: the fault is a lossy-sourced load, a load whose source register was last written by an entropy-losing op (or, mul, mulhi), whose saturated value recurs at (3/4)^32 per read and lands on one item through the era map; 96.6 percent of today's class v4 programs carry one; the bound H = W_0.1 from the 16 window draws (0.115 to 0.251 percent) plus the sum over load sites of h(last writer), h = 0.30 percent for or, 4.5 for an or chain, 0.067 mul, 0.049 mulhi, 0 for an injecting op or a rotate, with H at or under 1.2 x W_0.1, equal to the static rule "every load's source was last written by an injecting op or a rotate"; as a rejection it would cost 96.6 percent of candidates, so the form for v5 is a generator draw: a load's source drawn from the registers whose last writer injects, no attempts lost, rule (a) unchanged; its worth to a chip today, at most 1.005x on about half the hours, 1.048x on 5 percent, 1.067x at the ceiling; the project lead ruled option A at 15:2x UK by the coordinator's relay: the draw lands in class v4 itself (branch ca3-v4-amend, the hash lane, 0.3.19, a new program stream and vectors, the amended class with its own generator stamp), and class v5 inherits it by merging that branch into class-v5 once its commit is up, the pinned v5 pack re-exported on that base; this row is the record of the bound and its gate, satisfied by the v4 amendment; gate as above); beside it, the mixer-draw acceptance rule of the attack-pass lane's AP-F4-1 (`docs/analysis/attack-pass/f4-weakday.md` sections 6.7, 7 and 8 on branch attack-pass, relayed 10:0x UK): the day key's MUL block has a tail of cheap multipliers (low NAF weight), a tail of a sum with no weak class behind it; against the DSP-bound datapath F4 passes (0 of 2^28 days over 1.1x), on the LUT-adder metric 5,476 of 2^24 days are over 1.1x (the worst 1.173x; the worst calendar day in the first 100 years is chain day 29,337 at 1.121x), worth at most 12.1 percent more hash rate that day to a per-day LUT-recompute FPGA and nothing to a stored-dataset FPGA or any chip; the v5 rule for the day's mixer draw: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values, NAF weight at least 4 per word, at least 4 distinct ROT amounts, total rejection 6.1e-4 per day, the first calendar redraw at day 22,633 (5.2 years in), no devnet or testnet pack changes; gate: F4's harness (the 2^24-day census under both metrics, the worst day's M1 cost at or above 211 after the rule), re-gated by the attack-pass lane against this branch; it is the mixer-draw side of `memhard.rs` (`MixParams::with_shape`), independent of the leaves, and lands behind the v5 class; third, the attack-pass lane's AP-F1-1 shadow-block rule (`docs/analysis/attack-pass/f1-shadow.md`, harness `tools/attack/f1-shadow`, relayed 10:5x UK): over 100,000 class v4 programs the shadow block's peephole-removable instructions (a register written twice from one source with no write between: xor-cancel, sum-cancel, rotate merges, or-idempotence) average 0.62 percent of the 256 per pass, at most 5.078 percent (13 of 256, one program in 100,000 over 5 percent), nothing crossing a pass; clang -O3 removes the same from the honest kernel, so it is a bound on the shadow's useful work, not a chip shortcut; the v5 rule: the generator refuses a shadow block whose honest-compiler simplification exceeds a fraction and redraws from the next stream values; the fraction, set here from the census histogram (0.5-percent bins from 0: 55,595; 20,442; 11,790; 9,729; 1,447; 613; 256; 103; 17; 7; 1; 0): 3.0 percent, which rejects 384 of 100,000 draws (the bins from 3.0 up: 256 + 103 + 17 + 7 + 1, about 4e-3, under one attempt lost per 250 seeds, so the per-program spread and the acceptance rate of the 2.0 rule stay inside their bounds) and caps the removable share at 3 percent of the block where 5 percent would cap it at 5 for 1e-5 of draws; gate: F1's harness on 64 seeds of the v5 stream, every block under 3 percent, re-gated by the attack-pass lane; it is the shadow side of the v4 draw that v5 inherits from ca3-v4-amend and lands in `generator.rs` beside the shadow draw, behind the v5 class; the exec snapshot wire carrying the day streams (version 2); the GPU worker hosts' leaf buffer (`proto-cuda/nvrtc`, `proto-opencl`, `proto-metal`: the kernel text carries it, the hosts must upload it); the pool protocol's daily fetch (spec 09); the day-state witness in the pruning-proof format (with the class-signal witness); the spec text (01 1.8.5 the leaf line, 1.12 the cut and the reference block, 10 the witness); the 2019-class core row (O-1.14); Devnet 2 across a real day boundary. +Owed, named here so nobody looks for them: an acceptance-rule bound on a program's hot-set share (the Counter ASIC lane's item for v5, 09:5x UK, on main's word: the attack-pass lane's F8 found that on a class v4 program the top 0.1 percent of items take 0.52 percent of reads, 4.05x a uniform map, one item 153x the mean, most of it class v4's designed per-site windows of layer 8; the v5 rule rejects a draw whose windows and offsets coincide into a hot set beyond the window model's tail, with the rejection's cost stated in accepted programs per 64 seeds under the 2.0 bounds; the bound's definition and number come from the hash lane's window-model analysis on branch ca3-v4-uniform and the gate is F8's 64-seed census, `tools/attack/f8-uniform` on branch attack-pass phase E, the top 0.1 percent within the bound on every seed, re-gated by the attack-pass lane; nothing of it touches v4, which is on the vote; the bound landed 10:3x UK from `docs/analysis/ca3-v4-uniform.md` 095f84a7: the fault is a lossy-sourced load, a load whose source register was last written by an entropy-losing op (or, mul, mulhi), whose saturated value recurs at (3/4)^32 per read and lands on one item through the era map; 96.6 percent of today's class v4 programs carry one; the bound H = W_0.1 from the 16 window draws (0.115 to 0.251 percent) plus the sum over load sites of h(last writer), h = 0.30 percent for or, 4.5 for an or chain, 0.067 mul, 0.049 mulhi, 0 for an injecting op or a rotate, with H at or under 1.2 x W_0.1, equal to the static rule "every load's source was last written by an injecting op or a rotate"; as a rejection it would cost 96.6 percent of candidates, so the form for v5 is a generator draw: a load's source drawn from the registers whose last writer injects, no attempts lost, rule (a) unchanged; its worth to a chip today, at most 1.005x on about half the hours, 1.048x on 5 percent, 1.067x at the ceiling; the project lead ruled option A at 15:2x UK by the coordinator's relay: the draw lands in class v4 itself (branch ca3-v4-amend, the hash lane, 0.3.19, a new program stream and vectors, the amended class with its own generator stamp), and class v5 inherited it by merging that branch into class-v5 (4e737543, 13:0x UK). NOT satisfied yet (the coordinator's correction, 13:2x UK): the attack-pass lane's F8 re-gate on sub-version 1 has nine of the first 30 seeds over 1.2x (p31 at 29.3x) and 0.19 percent hot-set programs, because saturated values pass through saturation-preserving writers; class v5 draws under the amended v4 by merge and so inherits the gap, and takes the fix the same way when the Counter ASIC lane lands sub-version 2 inside generator 4 (byte 6 stays v5's); gate as above, open); beside it, the mixer-draw acceptance rule of the attack-pass lane's AP-F4-1 (`docs/analysis/attack-pass/f4-weakday.md` sections 6.7, 7 and 8 on branch attack-pass, relayed 10:0x UK): the day key's MUL block has a tail of cheap multipliers (low NAF weight), a tail of a sum with no weak class behind it; against the DSP-bound datapath F4 passes (0 of 2^28 days over 1.1x), on the LUT-adder metric 5,476 of 2^24 days are over 1.1x (the worst 1.173x; the worst calendar day in the first 100 years is chain day 29,337 at 1.121x), worth at most 12.1 percent more hash rate that day to a per-day LUT-recompute FPGA and nothing to a stored-dataset FPGA or any chip; the v5 rule for the day's mixer draw: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values, NAF weight at least 4 per word, at least 4 distinct ROT amounts, total rejection 6.1e-4 per day, the first calendar redraw at day 22,633 (5.2 years in), no devnet or testnet pack changes; gate: F4's harness (the 2^24-day census under both metrics, the worst day's M1 cost at or above 211 after the rule), re-gated by the attack-pass lane against this branch; it is the mixer-draw side of `memhard.rs` (`MixParams::with_shape`), independent of the leaves, and lands behind the v5 class; third, the attack-pass lane's AP-F1-1 shadow-block rule (`docs/analysis/attack-pass/f1-shadow.md`, harness `tools/attack/f1-shadow`, relayed 10:5x UK): over 100,000 class v4 programs the shadow block's peephole-removable instructions (a register written twice from one source with no write between: xor-cancel, sum-cancel, rotate merges, or-idempotence) average 0.62 percent of the 256 per pass, at most 5.078 percent (13 of 256, one program in 100,000 over 5 percent), nothing crossing a pass; clang -O3 removes the same from the honest kernel, so it is a bound on the shadow's useful work, not a chip shortcut; the v5 rule: the generator refuses a shadow block whose honest-compiler simplification exceeds a fraction and redraws from the next stream values; the fraction, set here from the census histogram (0.5-percent bins from 0: 55,595; 20,442; 11,790; 9,729; 1,447; 613; 256; 103; 17; 7; 1; 0): 3.0 percent, which rejects 384 of 100,000 draws (the bins from 3.0 up: 256 + 103 + 17 + 7 + 1, about 4e-3, under one attempt lost per 250 seeds, so the per-program spread and the acceptance rate of the 2.0 rule stay inside their bounds) and caps the removable share at 3 percent of the block where 5 percent would cap it at 5 for 1e-5 of draws; gate: F1's harness on 64 seeds of the v5 stream, every block under 3 percent, re-gated by the attack-pass lane; it is the shadow side of the v4 draw that v5 inherits from ca3-v4-amend and lands in `generator.rs` beside the shadow draw, behind the v5 class; the exec snapshot wire carrying the day streams (version 2); the GPU worker hosts' leaf buffer (`proto-cuda/nvrtc`, `proto-opencl`, `proto-metal`: the kernel text carries it, the hosts must upload it); the pool protocol's daily fetch (spec 09); the day-state witness in the pruning-proof format (with the class-signal witness); the spec text (01 1.8.5 the leaf line, 1.12 the cut and the reference block, 10 the witness); the 2019-class core row (O-1.14); Devnet 2 across a real day boundary. ## 12. The litepaper paragraph and the ledger row