Merge remote-tracking branch 'origin/master' into build-server
62
docs/analysis/ca3-v4-uniform.md
Normal file
|
|
@ -0,0 +1,62 @@
|
|||
# AP-F8-1 under the windows-union model: the hot set is the load source, not the window (7 October 2026)
|
||||
|
||||
Branch `ca3-v4-uniform` from master b92a5fd4, worker "v4-hash", on the attack-pass finding AP-F8-1 (`docs/analysis/attack-pass/f8-uniform.md`, branch attack-pass; logs `/srv/builds/igneum-wt-attack/target-attack-f8/log/`). Main's rulings bound this file: no generator change to class v4 on the live devnet; the analysis and its harness only. Every GPU-free number here is arithmetic on F8's logged counts or a run of the static census tool `tools/ca3-v4-uniform/` on igneum-build-1 (built through `tools/build-remote.sh`, rule R1); the chip figures are the terms of `docs/analysis/chip-model-v3.md` and are approximate.
|
||||
|
||||
## 1. The null F8's numbers must be read against
|
||||
|
||||
Layer 8 (`docs/plans/era-layout.md` 1.4, spec 01 1.13.1 as proposed) gives every load site a window draw `k_off = below(3)`: the site reads the whole dataset, an aligned half or an aligned quarter, at a 2^26-word floor. A quarter-window site concentrates its reads 4x on its quarter and a half-window site 2x on its half, by design; the union of the 16 windows is the whole dataset. For p1 (the devnet epoch-0 program, id c120d7963abdcd96) the 16 draws `0:2:1 1:1:1 2:1:1 3:1:1 4:0:0 5:1:1 6:0:0 7:2:2 8:1:1 9:1:1 10:2:0 11:0:0 12:0:0 13:0:0 14:2:0 15:1:1` (site:k:offset) give an expected read density by quarter of 3.25 : 2.25 : 5.75 : 4.75 sixteenths of the flat mean, which at 2^26 nonces (512 reads per item flat) is 416, 288, 736 and 608 reads per item. The top-f share of a Poisson mixture with those means (`uniform-model.txt`, exact Poisson for p1, a normal approximation for the census):
|
||||
|
||||
| Share of all reads on the top f of items, 2^26 nonces | Flat Poisson (F8's control) | Windows-union model, p1 | F8 measured, p1 | Beyond the window model |
|
||||
|---|---|---|---|---|
|
||||
| f = 0.1 percent | 0.115 | 0.160 | 0.520 | +0.36 |
|
||||
| f = 0.5 percent | 0.565 | 0.784 | 1.515 | +0.73 |
|
||||
| f = 1 percent | 1.120 | 1.553 | 2.495 | +0.94 |
|
||||
|
||||
So the window model moves the null from 0.115 to 0.160 percent at f = 0.1 percent (1.39x, not F8's 4.05x) and from 1.12 to 1.55 at f = 1 percent; it explains the 64-item-bucket sigma of p2 that F8 already attributed to the window layer, and it explains every per-site attribution row of p1 except one: sites with a half window over the hot region land 0.20 percent of their reads in the top 0.1 percent (sites 1, 2, 3, 5, 9 at 0.201), quarter-window sites 0 or about 0.4 (sites 0, 10, 14 at 0.000, site 7 at 0.241 straddling), whole-dataset sites 0.10. Site 15 lands 6.374 percent. The excess over the model (+0.36 at f = 0.1 percent) is one site.
|
||||
|
||||
## 2. The 153x item is the load source, and the model predicts it to the item
|
||||
|
||||
p1's site 15 is the load at instruction 63 (`load dst=3 src=6`, window half 1). Its source r6 was last written at instruction 61: `or dst=6 src=4` (`r6 |= r4`, `verify.rs` Op::Or), after a fresh dataset load into r6 at 47. An OR of two near-uniform registers sets each bit with probability 3/4, so the source takes the all-ones value with probability (3/4)^32 = 1.0e-4 per read and the values of popcount 31, 30, ... with 32, 496, ... times (3/4)^k (1/4)^(32 - k). The era map `y = rotl(x * 0x9ad30d99, 29)`, the half window and the interleave split (`memhard::Layout::split`, positions 0, 2, 12, 13) send x = 0xffffffff to item 0xca5b92: F8's hottest item exactly. F8's next seven items (0x8a5b92, 0xaa5b92, 0xba5b92, 0x825b92, 0xe65b92, 0x985b92, 0xbcc392) are exactly the seven one-zero-bit sources whose zero bit survives the window mask (bits 29, 28, 27, 26, 25, 24 and 15): 7 of 7. The measured count fixes the bit bias: 78,479 reads of 2^26 x 8 site-15 reads is p^32 at p = 0.7585 (r4 is slightly biased itself), and at that p the popcount model predicts 77,348 all-ones reads and 4.86 percent of site 15's reads into the top 0.1 percent of items (measured 6.37; 3.92 at p = 3/4). Per hash that is 4.86 / 16 = 0.30 percent of all reads, and 0.160 + 0.30 = 0.46 against F8's 0.520 at f = 0.1 percent; at f = 1 percent 14.5 / 16 = 0.90, and 1.55 + 0.90 = 2.46 against 2.495.
|
||||
|
||||
The same arithmetic for the other lossy writers (`uniform-model.txt`): a `mul` last writer zeroes the low bits by the operands' trailing zeros, so 1.07 percent of the site's reads land on the 0.1 percent of values with 10 or more trailing zeros (p2's `mul`-sourced sites 0 and 14 measured 0.971 and 0.966 percent); a `mulhi` last writer is dense near zero, 0.79 percent on the lowest 0.1 percent of values. An `or` whose operand was itself last written by `or` compounds the bias (3/4 to 7/8 to 15/16): p3's site 15 (`or` at 30, the load at 62) puts 72.4 percent of its reads into the top 0.1 percent, 4.6 percent of all reads on 16,777 items.
|
||||
|
||||
This is a fault class, not the window model: the acceptance rule's part (a) (`accept.rs` check_stale_loads) takes any write as a fresh source, and part (c)'s saturation count looks at the 16,384 final register values, not at a load's source mid-program, so an `or`, `mul` or `mulhi` as a load's last writer passes. The per-hash distinct-address check still holds (p1 127.999 items per hash; p3 127.97: a saturated site repeats its item inside a hash), and the acceptance rule's floor of 120 distinct of 128 admits exactly one site repeating its item in all 8 iterations and no more.
|
||||
|
||||
## 3. How common it is: the static census (`tools/ca3-v4-uniform`, 1,024 chain-shaped class v4 programs plus F8's p1 to p3)
|
||||
|
||||
For every load site, the op that last wrote its source in execution order (base instructions before it, else the shadow block of the previous iteration, else the base instructions after it): injecting (add, sub, xor, mad, shfl, load), bijective (rotl, rotr) or lossy (or, mul, mulhi). Run on igneum-build-1 (`uniform-census.txt`, binary sha256 ce9f83fe... then the narrowed chain rule).
|
||||
|
||||
| Census over 1,024 programs | Count | Share |
|
||||
|---|---|---|
|
||||
| Load sites by last writer: injecting / bijective / lossy | 11,368 / 2,121 / 2,943 of 16,432 | 69 / 13 / 18 percent; 2.87 lossy sites per program |
|
||||
| Programs with at least one lossy-sourced load | 992 | 96.6 percent |
|
||||
| ... with an `or`-sourced load (p1's class, 0.30 percent of all reads per site) | 498 | 48.5 percent |
|
||||
| ... with an `or`-of-`or` chain (p3's class, about 4.5 percent of all reads per site) | 50 | 4.9 percent |
|
||||
| ... with a `mul`-sourced load (0.067 percent per site) / a `mulhi`-sourced load (0.049) | 751 / 661 | 73.1 / 64.4 percent |
|
||||
| Predicted S_0.1 percent (window model plus the lossy sites): median / 90th / 99th / max | 0.45 / 0.88 / 5.29 / 9.82 percent | against the window model's 0.115 to 0.251 |
|
||||
| p1 / p2 / p3 predicted against F8 measured | 0.579 / 0.323 / 4.72 | 0.520 / 0.272 / 4.60 |
|
||||
|
||||
F8's proposed gate (the top 0.1 percent within 1.2x of the window-model control on every one of 64 seeds) fails 96.6 percent of today's programs, because any lossy-sourced site alone exceeds it (0.16 + 0.05 at the least); it is a generator change in a gate's clothing. A 2x bound fails 69.7 percent, 3x 48.9 percent; a bound of S_0.1 percent at or under 1 percent of all reads fails 6.9 percent (the `or` chains and the multi-`or` programs). The static rule "no load whose source's last writer is `or`" fails 48.4 percent; "no lossy last writer" 96.6 percent.
|
||||
|
||||
## 4. What the skew is worth to a chip (chip-model-v3.md terms, approximate)
|
||||
|
||||
A hot-set cache of the top 0.1 percent of items is 16,777 items x 64 B = 1.07 MB of SRAM, 0.53 mm^2 and $0.25 at 0.49 mm^2 and $0.23 per MB. It serves 0.52 percent of p1's reads (0.16 of them the window model's), 4.6 percent of p3's. The hash is latency-bound on its dependent reads, so a read served on die is time saved: a chip gains at most 1.005x on p1 and 1.048x on p3 from the cache. The ceiling under the live rule: part (c)'s 120-of-128 floor admits one site repeating its item in all 8 iterations and no more (two saturated sites fail it), so at most 8 of 128 reads, 6.25 percent, can sit on a constant item, and a chip's edge from this whole class is at most 1 / (1 - 0.0625) = 1.067x, in 64 bytes of SRAM, on the hours whose program carries such a site. The public claim rests on 2x margins (chip-model-v3.md); 1.067x does not move it, and the union of the windows is still the whole dataset every hour, so no window-level cache exists. What moves: per tier nothing in rate or watts (the honest card reads the hot item from L2 as the chip would), and the 5 percent rule of 2.0 is untouched.
|
||||
|
||||
## 5. The two options for the flip, priced (main's ruling 3; nothing ships on this without the project lead's word)
|
||||
|
||||
| Option | What changes | Cost | Risk |
|
||||
|---|---|---|---|
|
||||
| A. A class amendment in 0.3.19 before the flip: the generator draws a load's source from the registers whose last writer injects (or rule (a) tightened to the same), class v4 re-pinned | a new program stream: new vectors, the seven gate packs re-exported, the six gates again (the hash side G1 to G3 and the verifier re-run here in about an hour of Mac and PC 2 time; G4 to G6 the node lane), every node before the flip by the one-box-at-a-time fleet rule | hours of gate time, a fleet rollout, the 0.3.19 ship on the line | a node that misses the build splits the chain at the flip; the fix itself is small (one draw rule) |
|
||||
| B. Hold v4 at the floor as it is; the source rule in class v5 | nothing on the devnet; the attack-pass record carries the window null and the bound | a hot set on 48 percent of hours worth up to 1.005x to a chip, on 5 percent of hours up to 1.05x, 1.067x at the rule's ceiling, no chain risk | the public line must state the bound, not "uniform" |
|
||||
|
||||
The number that decides it: 1.067x at the ceiling against the 2x margin of the chip claim. Recommendation: B, with the v5 item below, unless the project lead wants the tail tight now.
|
||||
|
||||
## 6. The acceptance bound for the next class (main's ruling 4)
|
||||
|
||||
Definition: for a program, H = W_0.1(windows) + sum over load sites of h(last writer of the source), with W from the Poisson mixture of the 16 window draws (0.115 to 0.251 percent at 2^26 nonces) and h = 0.30 percent for `or`, 4.5 for an `or` chain, 0.067 for `mul`, 0.049 for `mulhi`, 0 for an injecting or bijective writer (the figures of section 2 at the measured bias). The bound: H at or under 1.2 x W, which is the static rule "every load's source was last written by an injecting op or a rotate" (any lossy writer breaks 1.2x). Its cost as a rejection rule on today's stream: 96.6 percent of candidates, about 30 attempts per seed on average. The cheaper form is a generator draw, not a rejection: draw a load's source from the registers whose last writer injects (today's rule draws from every written register), which costs no attempts and leaves rule (a) as it is. Either way the 64-seed census of F8's phase E is the gate, with the dynamic check extended to count saturated load sources over the 64 units beside the final values.
|
||||
|
||||
## 7. What is unverified
|
||||
|
||||
- The per-site h figures are the popcount and trailing-zeros models at the biases F8 measured on p1 and p2; p3's chain figure is F8's measurement, not a model. F8's phase E (64 seeds, dynamic) is the test of the whole table.
|
||||
- The window model's top-f shares for the census use a normal approximation per quarter (p1's exact Poisson 0.160 against 0.159).
|
||||
- No GPU run and no timing here; every number is a count or arithmetic.
|
||||
|
|
@ -41,7 +41,7 @@ Versions in the table: `igneum-pow` is the Rust crate at `igneum-pow/Cargo.toml`
|
|||
| 14 | Ethereum bytecode runs unchanged, with the documented differences of spec 7.1 | Homepage Build card; litepaper Building | tested by the team | as row 13; fixes `F-exec-A`, `F-exec-B` (spec 7.5) | `tools/evm-smoke/smoke.mjs`: deploy via viem, `increment`, `hashLoop`, `eth_estimateGas`, `eth_getLogs`; `tools/exec-attacks` scenarios 1 and 3; bench-log "execution layer attack fixes" | Deployment, calls, reverts, logs and gas estimates behave as viem expects; chain id 4463; the prototype pgas table gives 0.0095 to 0.028 pgas per gas, below the design's band before calibration, 3 October 2026. 4 October 2026: a transaction that would cross the block's proving budget is refused by the mempool and, if forced in, aborted and charged with its nonce advanced (25 of 25 checks; 30 of 30 malformed cases). Apple M5 Max. The `Prover` precompile, proof records and the shard planner are not in the node | none yet |
|
||||
| 15 | Every block is proven, with the proof landing within about a minute at launch | Homepage stats ("~60 s to a proof"); litepaper Proving; roadmap phase 3 gate | implemented | repo `d7e1f89` (GPU proof), `e01a3cc`, `292e800`, `eedd136` (`proving/igneum-prove`: shard cutter, MPT witnesses, shard and aggregator guests); SP1 6.8.1; spec 7.2, 7.6 | `proving/windows-wsl2` (SETUP-PROVER, PROVE-BLOCK) on the RTX 5090; `igneum-prove-host --mode block` on `proving/fixtures/`; bench-log "proving v0 on the RTX 5090" and "proving: devnet v4 shards" | First GPU proof of an Igneum block, 4 October 2026, RTX 5090 (WSL2, SP1 cuda, mining paused): fixture `block-78-increment` (2 transactions), core proof 1.4 s (7.3 MB, verify 0.221 s), compressed proof 2.7 s (1.27 MB, verify 0.038 s), post-state and receipts roots identical to the node's; 15.7x and 20.6x faster than a loaded M5 Max CPU. The same day on that CPU (load 38 to 47): a three-shard block proved shard by shard and aggregated by recursion, 19 min (1,139 s) end to end, 245 to 337 s per compressed shard proof, every proof verified. What is not there: no proof is produced, carried or checked on the chain (the devnet prover is a stub that signs claims), the proving pool pays nobody (row 21), the block proven is far below one shard, and the 60-second figure remains a design target; the pass mark is the standard in `docs/benchmarks/proving-e2e.md`. Second RTX 5090 run, 4 October 2026 evening (job run-20261004-173115): a full shard at the provisional S_p (6.75 M pgas, 60.8 M cycles) executed in 1.63 s, core proof 8.3 s (18.1 MB), compressed proof 10.9 s (1.27 MB, verify 0.040 s); a two-shard block (13.5 M pgas) proved shard by shard (11.7 s and 10.0 s) and aggregated in 2.2 s, 24 s of GPU stages end to end, every proof verified, six tampered witnesses rejected. The two host defects (an abort after the upload, an idle wait that turned out to be an unbuffered 18 MB proof save through the WSL2 file bridge, 24 minutes) are fixed (ledger P20) 5 October 2026, live devnet with real transactions (bench-log "real transactions, the first non-empty shard proven and paid"): block 72704 shard 0, 29 transfers, 5,800 pgas, proven on PC 2 in 34 s, verified on the Mac in 0.297 s and paid 1.7623 IGN, 53 s after the chain block executed; of about 1,400 blocks in the 20-minute window 36 were proven (the one prover takes the newest shard assigned to it), so "every block" is not yet true; a second content shard (72803, all copies skipped) failed the native-execution veto on the exporter's block structure, fixed with fixtures the same day, the node side pending the 0.3.9 rollout 5 October 2026, evening (bench-log "proving v1"): the aggregated segment record, the chain rule and the unproven rule are implemented behind `proving_v1_activation_daa` (branch proving-v1, not on the devnet before 0.3.11); on the RTX 5090 a chain of 8 consecutive live blocks proved and aggregated by recursion in 135.6 s with the miner on the card (17 s a block, one proof of 1,272,909 bytes attesting all 8, verified in 0.04 s); the 3-node fast-time harness paid a segment record 1.0 s after submission and refused a late one after its deadline (21 checks); the devnet itself, with one prover, carried proofs for 2.4% of blocks over 30 minutes at a block-to-record latency p50 44 s, p99 52 s. The "within about a minute" holds per proven block; "every block" needs 18 mining 5090s or 6 proving-only cards at empty blocks on the measured rates, and the mandatory rule stays off until the share is one | none yet |
|
||||
| 16 | A 12 GB card proves one shard in about 20 s (WITHDRAWN 5 October 2026: a 24 GB card proves a full shard at the adopted size in 4.3 s; 32 GB mines and proves) | Litepaper Proving ("The proving budget"); roadmap gate 2 | designed | spec 5.1 (Target), 7.6 (`S_p` provisional, 7,500,000 pgas = `B_p` / 4) | `PROVE-SHARD.bat` on the RTX 5090 (pending); the end-to-end standard in `docs/benchmarks/proving-e2e.md`; bench-log "proving: devnet v4 shards" | Measured on a 32 GB card, not yet on a 12 GB card. A shard at the provisional `S_p` is 60.8 M SP1 cycles on the prototype pgas table (9 cycles per pgas, 44 per EVM gas; the modexp entry about 100x its SP1 cost); on an RTX 5090 (4 October 2026 evening, job run-20261004-173115) it executed in 1.63 s and its compressed proof took 10.9 s, verified in 0.040 s, so the 32 GB card is inside the 20 s target with margin. Whether a 12 GB card proves it at all, and in what time, is the next measurement (an RTX 3060 and an RTX 5060 Ti 16 GB are on order). A per-shard time can be met by shrinking the shard, so the project does not use it as a pass mark 5 October 2026, evening (bench-log "proving v1", the S_p curve): measured on the RTX 5090 with SP1 6.8.1's GPU prover, the card to itself, 1-s nvidia-smi samples: an empty shard 13,874 MiB and 2.2 s; a full shard at the ADOPTED v1 budget (30,000 pgas, 4.7 M cycles) 20,434 MiB and 4.3 s; the full prototype shard (6.75 M pgas, 60 M cycles) 28,307 MiB and 10.8 s; beside the miner 15,670 and 30,039 MiB. No environment knob of SP1 moves the 13.9 GB floor and the GPU server has no options of its own, so on this build a 12 GB card proves nothing, a 16 GB card only empty shards, a 24 GB card the adopted full shard alone and beside the miner (22,210 MiB and 13.2 s, measured on the 32 GB card: the 5090's allocation pattern, not yet a run on a 24 GB card) and a 32 GB card the prototype shard beside the miner with 2.5 GB spare. The litepaper line now says so; the 12 GB gate returns when a prover build with a smaller floor is measured on a 12 GB card | none yet |
|
||||
| 17 | The chip resistance claim: the strongest chip in the public model reaches 5x to 9x per joule against an RTX 5090 today (modelled); class v4 brings it to 2.1x (k = 1) to 3.9x (k about 0.33, claimed by a withdrawn product) and its second rung to about 2.8x (modelled on measured watts); class v5 makes the dataset the chain's state so a stateless or stale chip is wrong on every item (designed, +0.2 ms verifier); the hot-set cache bounded at 1.067x at the ceiling (measured census of 1,024 programs) and the weak-day FPGA at most 12 percent on 12 days a century (measured census) are bounded and routed to the next class; datacentre silicon (H100 SXM, measured 7 October) does not change the question; a stored-dataset chip pays for itself only at about USD 100 M of market cap in two years (modelled) | The home page's chip line, the litepaper's chip model section, the miner page's line (the texts of `docs/plans/counter-asic-3-public-text-2026-10-07.md`) | tested by the team (every card, the verifier, the two attack-pass censuses, the H100), the chip itself modelled, class v5 and the ladder designed, the X9 core claimed and never measured | `docs/analysis/chip-model-v3.md` 5 and 6; `docs/analysis/latency-shadow-2026-10-06.md`; `docs/plans/counter-asic-3-status.md`; `docs/analysis/attack-pass/f8-uniform.md`, `f4-weakday.md`; `docs/design/class-v5-stored-state.md`; the H100 and market-cap rows of 7 October; `docs/plans/funding.md` (the three lots) | The chip model re-run on the measured class v3 and v4 rates, watts and verifier times; the hot-set census of 1,024 programs and the weak-day census of 2^24 days on the attack-pass branch; the H100 SXM bench row | 136 MH/s at 350 W (5090, bench) and 290 W (app); 27 MH/s at 21 W (M5 Max); 249 MH/s (H100 SXM) at 98 percent of its read ceiling, 1.78x hash, 1.15x MH/W, a third per rented dollar; 2.33 ms per warp; 5.1x to 9.2x; 2.1x, 3.9x, 2.8x; 1.067x at the ceiling; 12 percent on 12 days a century; 10.85 ms at rung 3; USD 100 M; 6 and 7 October 2026 | none yet; the three cryptanalysis lots are the next test |
|
||||
| 17 | The chip resistance claim: at launch the strongest chip in the public model reaches 2.1x (k = 1) to 3.9x (k about 0.33) per joule against an RTX 5090 under class v4, live from genesis on the testnet and the mainnet; the ladder's second rung brings it to about 2.8x; class v5 makes the dataset the chain's state so a stateless or stale chip is wrong on every item; the hot-set cache is bounded at 1.067x at the ceiling and the weak-day FPGA at 12 percent on 12 days a century, both routed to the next class; datacentre silicon does not change the question; a stored-dataset chip pays for itself only at about USD 100 M of market cap in two years; without class v4 the same chip would reach 5x to 9x (the class v3 baseline, the devnet's starting state, never the launch state) | the home page's chip line, the litepaper's chip section (/litepaper#chip-model), the miner page's line | tested by the team (every card, the verifier, the two attack-pass bounds, the H100), the chip itself modelled, class v5 and the ladder designed, the X9 core claimed and never measured | `docs/analysis/chip-model-v3.md` 5 and 6; `docs/analysis/latency-shadow-2026-10-06.md`; `docs/plans/counter-asic-3-status.md`; `docs/analysis/attack-pass/f8-uniform.md`, `f4-weakday.md`, `docs/analysis/ca3-v4-uniform.md`; `docs/design/class-v5-stored-state.md`; the H100 and market-cap rows of 7 October; `docs/plans/funding.md` (the three lots) | the chip model's arithmetic in its file; the card rows by the benchmark package; the attack-pass harnesses `tools/attack/f8-uniform` and the F4 census; the verifier by `igneum-pow bench` | 136 MH/s at 350 W (5090, bench) and 290 W (app); 27 MH/s at 21 W (M5 Max); 249 MH/s (H100 SXM) at 98 percent of its read ceiling, 1.78x hash, 1.15x MH/W, a third per rented dollar; 2.33 ms per warp; 2.1x, 3.9x, 2.8x at launch; 1.067x at the ceiling; 12 percent on 12 days a century; 10.85 ms at rung 3; USD 100 M; 5.1x to 9.2x the class v3 baseline; 6 and 7 October 2026, the M5 Max, PC 2's RTX 5090, PC 1's RX 9070 XT and RTX 4070, a rented H100 SXM, igneum-build-1 | none yet; the three cryptanalysis lots are the next test |
|
||||
| 18 | The chip resistance measurements: the program is latency-bound (random reads), not bandwidth-bound, on every card we own, and sits beyond a card's on-chip cache | Litepaper Mining ("waits on memory latency, not on maths or bandwidth"), vs RandomX; the numbers page | tested by the team | readwidth e752fc7 (`docs/plans/read-width.md`), ca2-era 78c0ee4, ca2-cache 2de19e5 (`docs/plans/hot-table.md`) | The dependent-read probes at 32 to 1,024 MiB and the hash rate per class on the three cards; the latency-bound share = rate over the probe ceiling per load | Latency-bound share at the 1 GiB dataset: RTX 5090 0.96 (v2) and 1.01 (v3), RX 9070 XT 0.87 and 0.95, M5 Max 1.01 and 1.06; wider reads do not close the AMD gap (the 9070 XT does 2.4 G dependent reads per second at every width; the 5090 goes bandwidth-bound at 64 B, share 0.58); a 32 to 96 MiB hot table is not kept resident by any card while the dataset streams (g 0.80 to 0.87 in the added form). 5 October 2026 | none yet |
|
||||
| 19 | The lottery hash is sound as a hash: uniform output, deterministic, no out-of-bounds read, fuzzed; class v3 bit-exact on the three vendors | Litepaper vs RandomX ("Every number above is measured and logged"), the numbers page | tested by the team | ca2-mixer 1ab8b21 (`tests/mixer.rs`, `tests/scratch.rs`), ca2-era 78c0ee4, ca2-soundness a465881 (`docs/analysis/scratch-soundness.md`), `igneum-pow/tests/packs.rs` | The crate suite (53 + 4 + 19 + 7), the Metal fuzz, edge, stats and determinism runs on the v3 construction, the pack vectors and 2^24 fingerprints on Metal, Apple OpenCL, the RTX 5090 and the RX 9070 XT, the 1,024-hash CPU re-check per card | Class v3 (mixer x8 + era): 200-program fuzz 200 of 200 on Metal, every tenth on Apple OpenCL; the pinned v3 packs 3/3 + 3/3 and 96 of 96 lanes on Metal and Apple OpenCL; the six era packs' fingerprints equal on the three vendors (PC 1 job run-ca2-era-pc1-20261005, 5 October 2026); the v2 exports byte-identical on the v3 crate; the final-class PC rows and the G2 re-check: job run-ca2-era-pc1b-20261005 (pending at the time of writing) | none yet |
|
||||
| 20 | No premine, no pre-sale, no allocation: every coin is minted by the schedule and every coin goes to the block producer (80%) and the proving pool (20%) | Homepage stats and Economics tiles; litepaper Supply, Economics | implemented | repo `6ac80a3`; fork "igneum-node devnet v0"; `consensus/core/src/igneum.rs`, `coinbase.rs` | `cargo test -p kaspa-consensus-core igneum` (8 pass: subsidy table, ramp, split, cap) and `cargo test -p kaspa-consensus coinbase` (8 pass); `igneum-miner inspect 40`; bench-log "igneum-node devnet v0" | Coinbases on the devnet: 80/20 exact on 39 of 39 single-payee blocks, the 20% to the `igneum-proving-pool-v0` output; the per-second schedule sums to under the 4,000,000,000 cap by less than 100 coins; 3,168,808,781 units per DAA second in years 0 to 2, halving at 63,115,200 DAA s. 3 October 2026, Apple M5 Max. The devnet genesis carries no allocation; the mainnet genesis does not exist yet, so the claim is about the code and the stated rule, not a launch that has happened | none yet |
|
||||
|
|
|
|||
|
|
@ -2435,6 +2435,35 @@ Same rule as the counts above. 167 entries.
|
|||
The bucket "Decided or closed by rule" counts each entry once: M8, M11 and P16 move there from Answered, so Answered is 28 less those three plus M16 and E17, as the table states; the sum of the rows is 167 with M8, M11, P16 and F3, F17 counted in Decided only.
|
||||
|
||||
|
||||
## Round 5 entry (7 October 2026, afternoon): the class v4 load-source amendment
|
||||
|
||||
### AP-F8-1. A load whose source was last written by `or`, `mul` or `mulhi` makes a cross-hash hot set
|
||||
"The item histogram of class v4 over 2^26 nonces is not uniform: the top 0.1 percent of items take 0.520 percent of reads against 0.115 for a uniform control (4.05x), one item takes 78,479 reads (153x the mean), and site 15 feeds 6.37 percent of its reads into that top 0.1 percent in every iteration." (attack-pass row F8, `docs/analysis/attack-pass/f8-uniform.md`, 7 October 2026)
|
||||
|
||||
Status: Fixed on a branch, pending the 0.3.20 node ship (7 October 2026, afternoon; the project lead's word 15:2x UK: option A, "do this but limit the testing, get it pushed"): `ca3-v4-amend` a0aaca92 (the generator and the packs) and 1748fd1d (the PC 2 playbook), on origin/master 5c08c6e0 plus `ca3-v4-uniform` 095f84a7 (the analysis). The node side rides `release-0.3.20-node`; the stamp is agreed with the node lane (a283f5f0d364ceef0). Was: a finding (7 October 2026, morning).
|
||||
|
||||
Answer: Correct as a fault, wrong as a null. The window layer (spec 01 1.13.1 as proposed, `docs/plans/era-layout.md` 1.4) moves the uniform null from 0.115 to 0.160 percent at the top 0.1 percent (1.39x, not 4.05x) and explains every per-site row of F8's attribution except site 15. Site 15's source r6 was last written by `or r6, r4` (instruction 61, the load at 63), a non-injective op whose output bits are 1 with probability 3/4, so the all-ones source recurs with probability (3/4)^32 per read; the era map sends it to item 0xca5b92, F8's hottest item exactly, and F8's next seven items are exactly the seven one-zero-bit sources whose zero survives the window mask. The popcount model at the measured bias (p = 0.7585) predicts 77,348 all-ones reads against 78,479, and the program's top-0.1-percent share at 0.58 against 0.52. The acceptance rule's part (a) takes any write as a fresh source and part (c) counts saturation on final register values only, so the class of fault passes it: of 1,024 chain-shaped class v4 programs 96.6 percent carry a load whose source's last writer is `or`, `mul` or `mulhi`, 48.5 percent an `or`-sourced one (0.30 percent of all reads per site), 4.9 percent an `or`-of-`or` chain (p3's class: 72 percent of that site's reads, 4.6 percent of all reads, on 0.1 percent of items). The ceiling under rule (c)'s 120-of-128 floor is one site repeating its item in all 8 iterations, 6.25 percent of reads, a chip edge of at most 1.067x in 64 bytes of SRAM; the public claim's 2x margin stands, and the public line says "bounded", not "uniform" (`docs/analysis/ca3-v4-uniform.md` sections 1 to 4).
|
||||
|
||||
The amendment (a0aaca92): in class v4's chain draw a load's source is drawn only from registers whose last writer injects (add, sub, xor, mad, shfl, load) or is a rotate (rotl, rotr), never one last written by `or`, `mul` or `mulhi` (`generator.rs` candidate_from_words_class, keyed on the era-composed V4_CLASS; no attempts lost, rule (a) unchanged; v2, v3 and the generator-2 ladder packs byte-identical). Split protection (the node lane's form): generator stays 4 and a generator-4 program id appends `"sub/" || PROGRAM_SUBVERSION_V4 (= 1) as little-endian u16` inside `program_id()`, so a binary from before the rule and one after it never share a program id for one seed and the node's id check catches a split; packs carry `IGNEUM_PROGRAM_SUBVERSION 1` and `"sub_version": 1`, and packcheck refuses a generator-4 pack whose sub-version is absent or other. The node side (release-0.3.20-node, built to a0aaca92): kaspa-pow's test pins 1a4230699a6b9c60 must-equal and c120d7963abdcd96 must-differ through `generator::program_id` with the suffix, and the amended class signals object byte 5 in the header (CLASS_SIGNAL_V4 = 5; the tally counts byte 5 and above), so a block of the 6 October stream (byte 4) never counts towards the flip and a byte-4 node forks alone at it; object 6 is class v5's. The seven gate packs re-exported (devnet epoch 0 and eras 0 to 5): id 1a4230699a6b9c60 (was c120d7963abdcd96, pinned as the must-differ vector in `tests/recheck.rs`), fingerprints Metal = Apple OpenCL 867dbc45cfb36b4d, 2146ecacc8c75a8e, fe52602393f6d3d4, 3b206471a13912b4, c3f03c4a5d7333aa, f1dfd7209f15bb97, 8c194da64fadf31d; the v3 control 73bcbfe8ccf988f1 / 90f794dd556f7a3b untouched; the zip of the eight packs sha256 889ec99976d2728b4b5035bfa476032e5b6a13b928968fc45236d5f25084aa39.
|
||||
|
||||
Evidence: `docs/analysis/ca3-v4-uniform.md` (the model, the census `tools/ca3-v4-uniform`, the chip pricing); F8's logs under `/srv/builds/igneum-wt-attack/target-attack-f8/log/`; the tests limited by the project lead's word to what prevents a split and proves the fix: the vectors (the seven packs' ids and fingerprints above), the `igneum-pow` crate suite on igneum-build-1 at 8c728ca3 (`tools/build-remote.sh -- test --release`, rc 0, 77 s, 12:1x UTC: 61 lib + 7 derive + 4 mixer + 19 packs + 2 recheck + 7 scratch = 100 passed, 0 failed, the pinned v2 and v3 packs byte-identical and the v4 must-equal and must-differ ids as pinned), the pairing of the fork's kaspa-pow with this igneum-pow on igneum-build-1 (PAIRING-LINE), and one G1 run on the RTX 5090 (PC 2 job run-ca3-v4-amend-g1-pc2-20261007, 09:41:07 to 09:41:28Z, exit 0, the installed 0.3.17 worker sha256 14b6637e..., NVRTC, beside the app's miner, the prover on, the lock held 09:40:22 to 09:41:51Z): self-test PASS on all eight packs and every 2^24 fingerprint equal to the Mac's (the seven above and the control 90f794dd556f7a3b), NVRTC 188 to 332 ms per pack, the 1 GiB build 38 to 49 ms.
|
||||
|
||||
Sub-version 2 (7 October 2026, afternoon; main's ruling B2: 0.3.20 ships sub-version 1 as object byte 5 untouched, and sub-version 2 is object byte 7, built on this branch at 07a809a7): the attack-pass lane's F8 census on sub-version 1 (64 chain-shaped seeds, the window-model control, 2^24 nonces) read 53 of 64 PASS and 11 FAIL at 1.2x, worst p31 at 29.27x, and named three residual classes, all a constant delivered through a writer the one-writer rule admits: saturation or zero preserved through rotl, rotr or a load after a saturated load (p6, p23, p26, p31, p34), zero from mulhi (p45), and the iteration boundary (an `or` at 63 feeding a load at 1, p11); p4, p10 and p25 at 1.28x to 1.57x are the window model's own tail. (An F9 hot-set census read on the same day was withdrawn by its author: its harness drew outside the rule and its metric counted the era's designed windows as hot; F8 is the one re-gate instrument.) Sub-version 2 closes the three: the draw takes a load's source only from a register fresh by dataflow (fresh at the start; a load keeps freshness only from a fresh source; add, sub, xor, mad, shfl from either operand; rotl, rotr from their operand; or, mul, mulhi never), keyed on the class v4 shape on every draw path; rule (a') of the acceptance rule runs that freshness to its fixpoint over the loop (base then shadow block) and rejects a candidate whose load reads an unfresh register in the steady state; rule (c') counts, per load site, the source values equal to 0 or all-ones over the 64 units' 16,384 evaluations and rejects at 164 or more. The devnet epoch-0 seed's attempt 0 is rejected and attempt 1 accepted: id a788661687db4bb3 (c120d7963abdcd96 and 1a4230699a6b9c60 the must-differ pair), fingerprints Metal = Apple OpenCL e370fb2080b7dbb1, b7237555d31fc3cf, b6b167fa15dfe2c9, 28bdf65eff33f2c4, e26d38c46f3f1b16, dd8fdf6ff4f59eed, 8bf40f5cb858d835 (13:01Z), the packs zip sha256 69c36772cd79e44e2ddd589466d9c64a94a13c9e970e9f27bd76feabb9b4581b; the seven 256-block ladder packs of `packs-ca3-shadow` re-exported under the rule (their measured rates stand as the old stream's). Nothing is proposed for sub-version 2 until the attack-pass lane's two gates on 07a809a7 are green (the 64-seed census under 1.2x on every seed, the hot-set census on the chain path); its suite, pairing (after the node lane's re-pin to byte 7) and G1 lines are added below as they land. The static census (`tools/ca3-v4-uniform`, box 2, 13:12Z) over 1,024 chain-shaped seeds plus F8's p1 to p3 at sub-version 2: 0 lossy-sourced load sites of 16,432 (14,329 injecting, 2,103 bijective), 0 programs with an `or`-, `mul`- or `mulhi`-sourced load, the no-era draw path giving the devnet epoch-0 seed the pack's own id a788661687db4bb3 (one stream on every path); the cost of rules (a') and (c'): 1.99 attempts per seed on average against 0.05 before (p2's seed took five), which is 2 ms of generation per rejected attempt on one core, nothing a miner or node notices. The crate suite at 526fa757 on box 2 (`tools/build-remote.sh --box 2 -- test --release`, route line "box 2 for class suite, priority normal", rc 0, 66 s, 13:15Z): 61 lib + 7 derive + 4 mixer + 19 packs + 2 recheck + 7 scratch = 100 passed, 0 failed, the pinned v2 and v3 packs byte-identical and the three v4 ids as pinned (a788661687db4bb3 equal; c120d7963abdcd96 and 1a4230699a6b9c60 differ).
|
||||
|
||||
AP-F8-2 (7 October 2026, 13:xx UTC, the attack-pass lane on sub-version 2 at 07a809a7): chain-shaped seed igneum-f9/331672 exhausted the 32 attempts under rule (a') and the generator panicked, which on the chain is an epoch no node can draw, a liveness halt; the measured (a') plus (c') rejection rate of about two thirds per attempt puts the exhaustion probability at about (2/3)^32, 2e-6 per epoch seed (sub-version 1 exhausted 0 of 10^6). Main's ruling: the draw is total and no consensus path panics. Fixed at 8bdcbdd8 with the stream unchanged (re-export diff 0; the id a788661687db4bb3 and the fingerprints stand, so sub-version 2 keeps its number and the running censuses): the attempt cap of the class v4 shape is 256 (`MAX_ATTEMPTS_V4`; v2 and v3 keep 32), which puts the exhaustion probability under 1e-45 at a worst-case draw of about half a second on one core; after the cap the seed takes the last-resort program, deterministic and accepted as drawn, the candidate at attempt 256 with every `or`, `mul` and `mulhi` of the base program and the shadow block rewritten to `xor`, so every register stays fresh from the init words on and rule (a') holds by construction. The spec text for class v4 therefore reads: attempts 0 to 255 under rules (a), (b), (a'), (c) and (c'), then the last-resort program; the probability of reaching it is (r)^256 for a per-attempt rejection rate r, under 1e-45 at the measured r of about 2/3. Tests: `class_v4_draw_is_total_with_the_last_resort` (the last resort on real (a')-rejected candidates, every load fresh after it, no lossy op left, the chain path over 64 seeds without a panic, the cap per class). The 10^6-seed exhaustion count at the fixed commit is the attack-pass lane's measurement (its re-gate string is 8bdcbdd8); the 4,096-seed census with the attempt histogram (box 2, 13:33Z, `tools/ca3-v4-uniform --n 4096 --f8` at 8bdcbdd8): 4,099 chain-shaped programs, 0 lossy-sourced load sites of 65,584 (57,304 injecting, 8,280 bijective), 0 exhaustions, the accepted attempt geometric with ratio 0.674 (1,338 at attempt 0, 921, 620, 408, 274, 189, 127, 75, 55, 40, 16, 11, 12, 6, 5, then one each at 16 and 17; mean 1.98), so the measured per-attempt rejection rate is r = 0.674 and the exhaustion probability is 0.674^32 = 3e-6 under the old cap and 0.674^256 = 1e-44 under the class v4 cap. The crate suite at 8bdcbdd8 on box 2 (route line "box 2 for class suite, priority normal", rc 0, 80 s, 13:31Z): 62 lib + 7 derive + 4 mixer + 19 packs + 2 recheck + 7 scratch = 101 passed, 0 failed, the total-draw test included. G1 for sub-version 2 on a fleet RTX 5090 (main's order; the fleet lane, p1-5090, driver 580.173.02, sm_120, the box's Linux NVRTC worker `igneum-worker-cuda 1.0 (4 October 2026)` sha256 97e036e2..., which carries no program and compiles each pack's own text; the kit zip sha256 asserted before the put; 13:43:22 to 13:44:07Z): `--bench --batches 5 --batch-log2 24 --block-warps 1` on all eight packs, self-test PASS on every pack with the cache FNV 448274a57f508cbc and every 2^24 fingerprint equal to the Mac's (the control 90f794dd556f7a3b and the seven above); the rates (120 to 142 MH/s beside the box's own miner loop) are a reference only. The Windows G1 on PC 2 (job run-ca3-v4-sub2-g1-pc2-20261007b, published 13:46:33Z under the PC 2 lock, ran 13:46:36 to 13:46:50Z, exit 0; app 0.3.19, the installed worker sha256 14b6637e..., the RTX 5090 switched off by the runner's --cards-off before the script and restored on exit, igneum-worker-cuda running 0 before and after, the prover untouched): self-test PASS on all eight packs with the cache FNV 448274a57f508cbc, every 2^24 fingerprint equal to the Mac's and the fleet's (the control 90f794dd556f7a3b and the seven above), NVRTC 197 to 317 ms per pack, the 1 GiB build 22 to 34 ms, 115 to 130 MH/s with the card alone (a reference, 5 batches).
|
||||
|
||||
AP-F8-2, the exhaustion half, FIXED-AND-PASSED at 8bdcbdd8 on the attack-pass lane's 10^6 chain-shaped seeds through the chain path (14:03:53Z): 0 exhausted, 0 panics, 4 seeds past attempt 31 (three at 32, one at 35; 4e-6, inside (2/3)^32), max attempt 35, no seed at the last resort; r = 0.67, mean 2.0 attempts per seed.
|
||||
|
||||
Sub-version 2's hot-set half did not read green: F8's 64-seed gate at 39 of 64 had 8 over 1.2x of the window model (p23 4.82x, p19 3.32x, p15 2.57x, p18 2.50x, p34 1.25x; p4, p8, p10 unattributed at 1.22x to 1.50x). AP-F8-3 (7 October 2026, 14:0x UTC, the hash lane): the cause of the whole residual is that `accept.rs` never executed the latency-shadow block. Its interpreter (`run_unit`) was written for class v2 and v3 and ran the 64 base instructions per iteration and nothing after instruction 63, while the hash (`verify.rs`, the kernels) runs the shadow 27 times at the end of every iteration; so every dynamic acceptance test (c), (c') judged a class v4 program the chain never hashes. Main's word (14:1x UTC): 0.3.21 ships object byte 5 (sub-version 1); sub-version 3 is 0.3.22's and starts with this fix. Sub-version 3, first commit: `run_unit` executes the shadow block after instruction 63 of every iteration, `reps` times with the iteration's sel, as the hash does; the test `acceptance_executes_the_shadow_block_as_the_verifier_does` pins the acceptance's execution to `verify.rs` on the devnet epoch-0 program and the six test eras (the output bit counts over the 64 units equal, and different with the shadow stripped), so the two paths cannot diverge silently again; `PROGRAM_SUBVERSION_V4` = 3 (a new acceptance verdict is a new stream); the devnet epoch-0 seed still accepts at attempt 1, so its program and fingerprint are sub-version 2's (e370fb2080b7dbb1) under the new id a785001687d8688a (the must-differ set: c120d7963abdcd96, 1a4230699a6b9c60, a788661687db4bb3); the packs zip sha256 4f2445c50c58d76a5544023492d8b858d0b07c5e372d31f9c90c4ce51f829154. The class behind p23, localised from its program and reproduced in the acceptance's own execution: site 7 (instruction 38) reads r6 after 25 `mulhi r6`, 31 `or r6 |= r4`, 35 `xor r6 ^= r4`, which is `r6 & ~r4`, an AND mask the lineage rule counted as fresh because the xor's operand is the or's; over 2^20 evaluations on the closed-form words site 7 reads 874,953 distinct word indices against about 1,046,500 for every other site (0.84 of uniform; 2.2 s on one core), over 2^24 8,979,203 against about 16,260,000 (0.55; 35 s). The second sub-version 3 commit (held, prepared in the worktree) is a per-site distinct-index ratio against the uniform expectation of the site's window, its threshold set from the clean seeds' spread and its sample size from the cost line above; the dynamic bounds as first specified (a most-repeated-value bound at 16,384 and a distinct floor at 2^19.5 over 2^20) do not reach p23 and are not committed. The crate suite at ddacfbd3 on box 2 (route "box 2 for class suite, priority normal", rc 0, 41 s, 14:20Z): 63 lib + 7 derive + 4 mixer + 19 packs + 2 recheck + 7 scratch = 102 passed, 0 failed, the agreement test included. F8's final 64-seed table on sub-version 2 (the attack-pass lane, 14:2x UTC): 9 over 1.2x (p23 4.82x, p19 3.32x, p15 2.57x, p18 2.50x, p56 2.01x, p10 1.50x, p8 1.38x, p34 1.25x, p4 1.22x), the 55 clean seeds at 0.9915x to 1.144x; the 256-item bucket entropy over the window separates the strong four only (0.637 to 0.974 against a clean minimum of 0.9865 over 848 site rows), so the threshold of the second commit's distinct-index ratio comes from a run of that ratio on the 55 clean seeds at 2^20. The static census at ddacfbd3 (box 2, 14:22Z, 4,096 chain-shaped seeds plus F8's p1 to p3): 4,099 programs, 0 lossy-sourced load sites of 65,584, 0 exhaustions, the accepted attempt geometric as before (1,328 at attempt 0, 917, 622, 409, 284, ... one each at 16 and 17; mean 1.998, max 17), so the shadow-executed verdicts move a handful of seeds' attempts and nothing else; the devnet epoch-0 seed at attempt 1, id a785001687d8688a.
|
||||
|
||||
Sub-version 3, second commit (7 October 2026, 14:44 UTC, the hash lane; the sub-version number stays 3 and the stream is unchanged: re-export diff 0 on the eight packs, the id a785001687d8688a, the seven fingerprints and the zip sha256 stand), two rules. The shared-operand rule, in the draw's source rule and in the acceptance's (a') pass: or-then-xor or or-then-sub on one operand is `d & ~s`, xor-then-or is `d | s`, so the second write leaves the register lossy although either op alone injects; any write to either register clears the relation. On p23 the chain's attempt 1 (id d65122675f16a1c7) now draws site 7 (instruction 38) from r5 and passes; the devnet epoch-0 seed still draws attempt 1, the same program. Rule (c''), the distinct-index ratio: over 4,096 units (2^20 evaluations per site, the closed-form words, the shadow executed) every load site's count of distinct word indices against the uniform expectation on its window (N - N^2 / 2W, the window 2^28 >> min(win, 2)) must reach 0.98 (`MIN_DISTINCT_RATIO_V4`), the last test of the chosen candidate; a candidate under it is rejected and the next attempt drawn under the 256 cap and the last resort. The floor from the 64-seed run at 2^20 (box 2): the 55 clean seeds' minimum over their site rows 0.9960 (p1 0.9990, median 1.0000); the strong five p23 0.8361, p18 0.9274, p19 0.9335, p15 0.9432, p56 0.9654; 0.98 sits 0.015 from each side. The chain's own candidates it refuses (the test `class_v4_distinct_ratio_rejects_the_low_entropy_band`, box 2, 2.1 to 2.2 s each): p15 attempt 3 id 52638ea2e8b0fd68 site 2 at 0.943, p18 attempt 2 id 9a37e9489d8ba698 site 6 at 0.927, p19 attempt 0 id 79d7441de0689223 site 15 at 0.933, p56 attempt 2 id 486a8ad2701ec3b5 site 2 at 0.965; p23's attempt 1 with site 7 put back to r6 is refused by (a') (`UnfreshLoadSource`) and, run anyway, by the ratio at 0.836 (874,928 distinct of 1,048,576). Main's rule for a staged 2^24 pass (taken only if 2^24 separates the weak four from the clean seeds by at least the 2^20 gap) was decided by the 2^24 lines (box 2, 35 s per seed on one core): the weak four p34 0.9181, p4 0.9614, p8 0.9630, p10 0.9612; the clean seeds p44 0.9612, p52 0.9613, p3 0.9971, p2 and p5 1.0004; p23's attempt 1 1.0004. Two clean seeds sit on the weak four's value, so a 2^24 floor that reaches the weak four rejects clean seeds, and the 0.961 that recurs on both sides is a band the ratio reads at 2^24 that F8's hot-set gate did not flag on p44 or p52. Committed: the ratio at 0.98 over 2^20 alone (`ACCEPT_UNITS_DISTINCT_V4` = 4096), no 2^24 stage. Open tail: p4, p8, p10 and p34 (1.22x to 1.50x on F8's gate) read 0.9927 to 0.9963 at 2^20, inside the clean spread; unattributed and chased. The (B) most-repeated-value bound stays in the file unwired (`MAX_SOURCE_REPEAT_V4`, `most_repeated`). Cost: the chosen candidate's acceptance gains one 2^20 pass, 2.1 to 2.2 s on one box-2 core, once per epoch draw per node.
|
||||
|
||||
The second sub-version 3 commit is 017e70376489251e18564c0abce7e466e606c8b3 (pushed 15:10 UTC's preceding hour, pre-push gate GREEN, CI run 37639406567 success). The crate suite at 017e7037 on box 2 (rc 0, 184 s): 64 lib + 7 derive + 4 mixer + 19 packs + 2 recheck + 7 scratch = 103 passed, 0 failed, 3 diagnostics ignored; the lib tests took 140.8 s against ddacfbd3's 41 s for the whole suite, because every class v4 draw in the tests now pays the 2^20 pass on its chosen candidate (2.2 s each); that is CI time, not node time. The static census at 017e7037 (box 2, 15:1x UTC, 4,096 chain-shaped seeds plus F8's p1 to p3, the draws in parallel over the box's cores since the ratio pass makes the serial run a four-hour job): 4,099 programs, 0 lossy-sourced load sites of 65,584 (57,322 inject, 8,262 bijective), 0 exhaustions; the accepted attempt 1,297 at attempt 0, 899, 609, 416, 298, 197, 125, 92, 54, 46, 21, 14, 13, 9, 6, 0, 2, 1 (mean 2.086, max 17) against ddacfbd3's 1,328, 917, 622, 409, 284, ... (mean 1.998, max 17): the ratio refuses about 4 percent of the candidates that pass every other test, one more attempt on about one seed in twelve; the devnet epoch-0 seed at attempt 1, id a785001687d8688a; F8's p2 at attempt 5, p3 at attempt 1. The attack-pass lane's class check of ddacfbd3's shadow-executed verdicts against 8bdcbdd8 (598,678 chain-shaped seeds): 11,990 (2.0 percent) accept at a different attempt, 0 exhausted, max attempt 32; on 017e7037 the pairing holds (the devnet epoch-0 program draws as a785001687d8688a) and its 64-seed gate at 2^24 and 10^6 exhaustion count were running at 15:1x UTC (finish about 16:05 UTC).
|
||||
|
||||
F8's 64-seed gate on 017e7037 (the attack-pass lane, box 2, last seed 16:00:20 UTC): 60 of 64 under 1.2x; the four over are the named tail and nothing else (p10 1.5036x, p8 1.3776x, p34 1.2505x, p4 1.2167x; hottest items 355 to 541 reads of 2^31, unattributed, inside the ratio's clean spread); p23, p19, p15, p18 and p56 under the line; the clean spread 0.9915x to 1.144x. Exhaustion on 017e7037: 0 in the attack-pass lane's 20,532 chain-shaped seeds (max attempt 29) plus this lane's 4,099 (max 17); the 10^6 count continues as a strengthening line. Cost line for the node: the chain's epoch draw now pays the 2^20 ratio pass on every candidate that reaches it (the accepted one, and the about 4 percent that fail there), 2.2 s per pass on one box-2 core, so about 2.3 s per epoch draw on that core and, by the attack-pass lane's reading, about 4 to 5 s per epoch per node on slower cores; the earlier rejections cost milliseconds. From the attack-pass lane sub-version 3 at 017e7037 reads green on both gates with the named tail; its close line went to the plan, main and the Counter ASIC lane. This row stays open on the tail (p4, p8, p10, p34) until it is attributed or ruled accepted.
|
||||
|
||||
Owed (recorded, not run, by the project lead's word): G2 (the CPU verifier on 1,024 hashes per card) on the amended stream; G3 (the Metal fuzz, edge, stats and determinism runs) on the amended stream; the hash-rate ladder re-measure on the M5 Max and the RTX 5090 (the amendment changes the base program's source draws, not the op mix or the load count, so the latency-bound rows of `docs/analysis/latency-shadow-2026-10-06.md` are expected to hold within their spread; unmeasured); AMD (the RX 9070 XT, PC 1); the 2019-class verifier core (O-1.14); F8's phase E (the 64-seed dynamic census) on the amended stream, which is the attack-pass lane's and the test of the per-op table. The row reads FIXED-AND-PASSED only after phase E passes against the amended class.
|
||||
|
||||
## Genesis forward-compatibility entries (7 October 2026, mission item 8, branch `genesis-forward`)
|
||||
|
||||
The three genesis fields of `docs/analysis/mission/mission.md` section 2.8, built on the node fork branch `genesis-forward` (from release-0.3.19-node dc141409) and the repo branch `genesis-forward`; the design and the gates in `docs/design/genesis-forward.md`. Every switch is never on the devnet (its digest c562d70e... does not move); the testnet genesis sets all three (the testnet lane re-pins and re-digests).
|
||||
|
|
|
|||
|
Before Width: | Height: | Size: 224 KiB After Width: | Height: | Size: 215 KiB |
|
Before Width: | Height: | Size: 225 KiB |
|
Before Width: | Height: | Size: 238 KiB After Width: | Height: | Size: 222 KiB |
|
Before Width: | Height: | Size: 240 KiB |
BIN
docs/plans/site-ui-5-shots/app-band__1440-dark.png
Normal file
|
After Width: | Height: | Size: 31 KiB |
|
Before Width: | Height: | Size: 1.1 MiB After Width: | Height: | Size: 1 MiB |
|
Before Width: | Height: | Size: 1.2 MiB |
|
Before Width: | Height: | Size: 586 KiB After Width: | Height: | Size: 553 KiB |
|
Before Width: | Height: | Size: 597 KiB |
|
Before Width: | Height: | Size: 478 KiB After Width: | Height: | Size: 476 KiB |
|
Before Width: | Height: | Size: 478 KiB |
|
Before Width: | Height: | Size: 367 KiB After Width: | Height: | Size: 374 KiB |
|
Before Width: | Height: | Size: 364 KiB |
|
Before Width: | Height: | Size: 549 KiB After Width: | Height: | Size: 537 KiB |
|
Before Width: | Height: | Size: 548 KiB |
|
Before Width: | Height: | Size: 549 KiB After Width: | Height: | Size: 536 KiB |
|
Before Width: | Height: | Size: 547 KiB |
|
Before Width: | Height: | Size: 332 KiB After Width: | Height: | Size: 312 KiB |
|
Before Width: | Height: | Size: 332 KiB |
|
Before Width: | Height: | Size: 277 KiB After Width: | Height: | Size: 262 KiB |
|
Before Width: | Height: | Size: 276 KiB |
|
Before Width: | Height: | Size: 590 KiB After Width: | Height: | Size: 567 KiB |
|
Before Width: | Height: | Size: 605 KiB |
|
Before Width: | Height: | Size: 518 KiB After Width: | Height: | Size: 492 KiB |
|
Before Width: | Height: | Size: 528 KiB |
|
Before Width: | Height: | Size: 517 KiB After Width: | Height: | Size: 475 KiB |
|
Before Width: | Height: | Size: 516 KiB |
|
Before Width: | Height: | Size: 309 KiB After Width: | Height: | Size: 249 KiB |
|
Before Width: | Height: | Size: 311 KiB |
|
Before Width: | Height: | Size: 381 KiB After Width: | Height: | Size: 350 KiB |
|
Before Width: | Height: | Size: 385 KiB |
|
Before Width: | Height: | Size: 263 KiB After Width: | Height: | Size: 255 KiB |
|
Before Width: | Height: | Size: 266 KiB |
|
Before Width: | Height: | Size: 335 KiB After Width: | Height: | Size: 300 KiB |
|
Before Width: | Height: | Size: 335 KiB |
|
Before Width: | Height: | Size: 276 KiB After Width: | Height: | Size: 255 KiB |
|
Before Width: | Height: | Size: 276 KiB |
|
Before Width: | Height: | Size: 267 KiB |
|
Before Width: | Height: | Size: 82 KiB |
|
Before Width: | Height: | Size: 844 KiB After Width: | Height: | Size: 1.1 MiB |
|
Before Width: | Height: | Size: 1.2 MiB |
|
Before Width: | Height: | Size: 558 KiB After Width: | Height: | Size: 571 KiB |
|
Before Width: | Height: | Size: 586 KiB |
|
Before Width: | Height: | Size: 328 KiB After Width: | Height: | Size: 304 KiB |
|
Before Width: | Height: | Size: 336 KiB |
|
Before Width: | Height: | Size: 327 KiB After Width: | Height: | Size: 305 KiB |
|
Before Width: | Height: | Size: 337 KiB |
|
Before Width: | Height: | Size: 345 KiB After Width: | Height: | Size: 344 KiB |
|
Before Width: | Height: | Size: 346 KiB |
|
Before Width: | Height: | Size: 374 KiB After Width: | Height: | Size: 350 KiB |
|
Before Width: | Height: | Size: 372 KiB |
|
Before Width: | Height: | Size: 488 KiB After Width: | Height: | Size: 480 KiB |
|
Before Width: | Height: | Size: 489 KiB |
|
Before Width: | Height: | Size: 396 KiB After Width: | Height: | Size: 381 KiB |
|
Before Width: | Height: | Size: 398 KiB |
|
Before Width: | Height: | Size: 414 KiB After Width: | Height: | Size: 504 KiB |
|
Before Width: | Height: | Size: 611 KiB |
|
Before Width: | Height: | Size: 378 KiB After Width: | Height: | Size: 368 KiB |
|
Before Width: | Height: | Size: 386 KiB |
|
Before Width: | Height: | Size: 407 KiB After Width: | Height: | Size: 377 KiB |
|
Before Width: | Height: | Size: 407 KiB |
|
Before Width: | Height: | Size: 350 KiB After Width: | Height: | Size: 340 KiB |
|
Before Width: | Height: | Size: 351 KiB |
BIN
docs/plans/site-ui-5-shots/miner-band__1440-dark.png
Normal file
|
After Width: | Height: | Size: 40 KiB |
|
Before Width: | Height: | Size: 571 KiB After Width: | Height: | Size: 552 KiB |
|
Before Width: | Height: | Size: 682 KiB |
|
Before Width: | Height: | Size: 402 KiB After Width: | Height: | Size: 401 KiB |
|
Before Width: | Height: | Size: 285 KiB |
|
Before Width: | Height: | Size: 225 KiB After Width: | Height: | Size: 223 KiB |
|
Before Width: | Height: | Size: 223 KiB |
|
Before Width: | Height: | Size: 244 KiB After Width: | Height: | Size: 213 KiB |
|
Before Width: | Height: | Size: 242 KiB |
|
Before Width: | Height: | Size: 426 KiB After Width: | Height: | Size: 417 KiB |
|
Before Width: | Height: | Size: 425 KiB |
|
Before Width: | Height: | Size: 333 KiB After Width: | Height: | Size: 321 KiB |
|
Before Width: | Height: | Size: 331 KiB |
|
Before Width: | Height: | Size: 396 KiB After Width: | Height: | Size: 401 KiB |
|
Before Width: | Height: | Size: 393 KiB |
|
Before Width: | Height: | Size: 396 KiB After Width: | Height: | Size: 411 KiB |
|
Before Width: | Height: | Size: 392 KiB |
BIN
docs/plans/site-ui-5-shots/wallet-band__1440-dark.png
Normal file
|
After Width: | Height: | Size: 30 KiB |
BIN
docs/plans/site-ui-5-shots/wallet-band__390-dark.png
Normal file
|
After Width: | Height: | Size: 26 KiB |
BIN
docs/plans/site-ui-5-shots/wallet-full__1440-dark.png
Normal file
|
After Width: | Height: | Size: 1.7 MiB |
BIN
docs/plans/site-ui-5-shots/wallet-full__390-dark.png
Normal file
|
After Width: | Height: | Size: 1.2 MiB |
|
Before Width: | Height: | Size: 676 KiB After Width: | Height: | Size: 652 KiB |
|
Before Width: | Height: | Size: 201 KiB |
|
Before Width: | Height: | Size: 429 KiB After Width: | Height: | Size: 411 KiB |
|
Before Width: | Height: | Size: 451 KiB |
|
|
@ -15,13 +15,31 @@
|
|||
//! seed words 2 and 3 at its multiply-shift index, a second pure table beside the dataset stand-in (words 0 and 1). The census (section 7.3) checked on 100,000 programs that the
|
||||
//! closed-form verdict agrees with the memory-hard one on all but 39 threshold-edge cases.
|
||||
|
||||
use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES};
|
||||
use crate::generator::{Instr, LoadClass, Op, Program, ShadowClass, INSTR_COUNT, ITERATIONS, LANES, V4_CLASS, V4_SHADOW_INSTRS};
|
||||
use crate::seed::{fnv1a64, SplitMix64};
|
||||
use crate::memhard::hot_index;
|
||||
use crate::verify::{dataset_elem, fold_words, load_index, splitmix32, ScratchModel};
|
||||
|
||||
/// Units (32-lane warps) the dynamic test interprets.
|
||||
pub const ACCEPT_UNITS: usize = 64;
|
||||
|
||||
/// Class v4 sub-version 3, rule (c''): the per-site distinct-index RATIO (AP-F8-1's low-entropy-band class, 7 October
|
||||
/// 2026, the attack-pass gate's numbers through main), keyed on the class v4 shape. Over [`ACCEPT_UNITS_DISTINCT_V4`]
|
||||
/// units (2^20 evaluations per site) the count of distinct dataset word indices a load site reads, against the
|
||||
/// expectation of a uniform source on the site's window (N - N^2 / 2W), must reach [`MIN_DISTINCT_RATIO_V4`]. The
|
||||
/// floor sits between the 55 clean F8 seeds' minimum over their site rows (0.9960; the clean p1 0.9990, the median
|
||||
/// 1.0000) and the strong failing seeds' maximum (p56 0.9654; p23 0.8361, p18 0.9274, p19 0.9335, p15 0.9432),
|
||||
/// 0.015 from each. The open tail: F8's p4, p8, p10 and p34 (1.22x to 1.50x on the gate) read 0.9927 to 0.9963 at
|
||||
/// 2^20, inside the clean spread, and a 2^24 pass does not separate them either (p34 0.9181, p4 0.9614, p8 0.9630,
|
||||
/// p10 0.9612 against the clean p44 0.9612, p52 0.9613, p3 0.9971, p2 and p5 1.0004); they stay unattributed and
|
||||
/// chased in `docs/fud-ledger.md` AP-F8-1. A candidate under the floor is rejected and the next attempt drawn under
|
||||
/// the 256 cap and the last resort. Measured on one box-2 core with the shadow executed: 2.8 s per chosen candidate.
|
||||
pub const ACCEPT_UNITS_DISTINCT_V4: usize = 4096;
|
||||
/// The ratio floor at 2^20.
|
||||
pub const MIN_DISTINCT_RATIO_V4: f64 = 0.98;
|
||||
/// Kept for the record and the driver, not wired: the most repeated source value per site over the (c) units'
|
||||
/// 16,384 evaluations (a uniform site repeats a value 2 or 3 times; the finding's bands sit under the ratio instead).
|
||||
pub const MAX_SOURCE_REPEAT_V4: u32 = 8;
|
||||
/// Hashes the dynamic test evaluates: 2,048.
|
||||
pub const ACCEPT_HASHES: usize = ACCEPT_UNITS * LANES;
|
||||
/// Domain tag of the base-nonce stream.
|
||||
|
|
@ -57,6 +75,19 @@ pub enum Reject {
|
|||
LaneConstantSite { iteration: u8, instr: u8, unit: u8 },
|
||||
/// (c): `count` final register values were 0 or all ones.
|
||||
Saturated { count: u32 },
|
||||
/// (a'), class v4 sub-version 2 (AP-F8-1): the load at `instr` reads `reg`, which is not fresh by dataflow in the
|
||||
/// steady state of the loop (the freshness fixpoint over the base program and the shadow block).
|
||||
UnfreshLoadSource { instr: u8, reg: u8 },
|
||||
/// (c'), class v4 sub-version 2 (AP-F8-1): the load at `site` read a source value of 0 or all ones in `count` of
|
||||
/// its 16,384 evaluations (64 units x 32 lanes x 8 iterations); limit [`MAX_SATURATED`] - 1, the same 1 percent as (c).
|
||||
SaturatedSource { site: u8, count: u32 },
|
||||
/// (c'') (B), class v4 sub-version 3: the load at `site` read the value `value` in `count` of its 16,384 (c)
|
||||
/// evaluations (limit [`MAX_SOURCE_REPEAT_V4`] - 1): one constant upstream that the lineage rule cannot see.
|
||||
RepeatedSource { site: u8, value: u32, count: u32 },
|
||||
/// (c''), class v4 sub-version 3: the load at `site` read `distinct` distinct dataset word indices over
|
||||
/// `evaluations`, `ratio_milli` / 1000 of a uniform source on its window, under the floor: a low-entropy index band
|
||||
/// (F8's p23, p18, p19, p15, p56).
|
||||
LowEntropySite { site: u8, distinct: u32, evaluations: u32, ratio_milli: u32 },
|
||||
/// (c): output bit `bit` was set in `ones` of 2,048 hashes.
|
||||
OutputBias { bit: u8, ones: u32 },
|
||||
/// (c): the distinct-address sum was `sum`.
|
||||
|
|
@ -75,6 +106,10 @@ impl std::fmt::Display for Reject {
|
|||
write!(f, "(c) load at iteration {iteration} instruction {instr} reads one address in all lanes of unit {unit}")
|
||||
}
|
||||
Reject::Saturated { count } => write!(f, "(c) {count} of 16384 final register values saturated (limit 163)"),
|
||||
Reject::UnfreshLoadSource { instr, reg } => write!(f, "(a') load at {instr} reads r{reg}, not fresh by dataflow in the loop's steady state (class v4 sub-version 2)"),
|
||||
Reject::RepeatedSource { site, value, count } => write!(f, "(c'') load site {site} read the value {value:#010x} in {count} of 16384 evaluations (limit {})", MAX_SOURCE_REPEAT_V4 - 1),
|
||||
Reject::LowEntropySite { site, distinct, evaluations, ratio_milli } => write!(f, "(c'') load site {site} read {distinct} distinct word indices over {evaluations} evaluations, {}.{:03} of a uniform source on its window (floor {MIN_DISTINCT_RATIO_V4} at 2^20)", ratio_milli / 1000, ratio_milli % 1000),
|
||||
Reject::SaturatedSource { site, count } => write!(f, "(c') load site {site} read a saturated source value in {count} of 16384 evaluations (limit 163)"),
|
||||
Reject::OutputBias { bit, ones } => write!(f, "(c) output bit {bit} set in {ones} of 2048 hashes"),
|
||||
Reject::DistinctAddresses { sum } => {
|
||||
write!(f, "(c) distinct dataset addresses {sum} over 2048 hashes (mean {:.2}, needs above 120 of 128 of the dataset loads)", *sum as f64 / 2048.0)
|
||||
|
|
@ -136,28 +171,170 @@ fn check_injecting_writes(instrs: &[Instr]) -> Result<(), Reject> {
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Parts (a) and (b).
|
||||
/// The most repeated value of `values` (sorted in place) and that value: (B), kept for the driver, not wired.
|
||||
#[allow(dead_code)]
|
||||
fn most_repeated(values: &mut [u32]) -> (u32, u32) {
|
||||
values.sort_unstable();
|
||||
let (mut best, mut best_v, mut run) = (0u32, 0u32, 0u32);
|
||||
for i in 0..values.len() {
|
||||
run = if i > 0 && values[i] == values[i - 1] { run + 1 } else { 1 };
|
||||
if run > best {
|
||||
best = run;
|
||||
best_v = values[i];
|
||||
}
|
||||
}
|
||||
(best, best_v)
|
||||
}
|
||||
|
||||
/// (c''), class v4 sub-version 3: the 2^20 ratio pass on the chosen candidate (the constants above).
|
||||
pub fn check_distinct_indices_v4(p: &Program) -> Result<(), Reject> {
|
||||
distinct_ratio_pass(p, ACCEPT_UNITS_DISTINCT_V4, MIN_DISTINCT_RATIO_V4).map(|_| ())
|
||||
}
|
||||
|
||||
/// One ratio pass over `units`: every load site's distinct word indices against the uniform expectation on its
|
||||
/// window (`N - N^2 / 2W`, the window `2^28 >> min(win, 2)` words of the closed-form dataset), `Err` at the first
|
||||
/// site under `floor`, else the minimum ratio and its site.
|
||||
pub fn distinct_ratio_pass(p: &Program, units: usize, floor: f64) -> Result<(f64, usize), Reject> {
|
||||
let n = (units * LANES * ITERATIONS) as f64;
|
||||
let d = distinct_indices_v4(p, units)?;
|
||||
let mut min = (f64::MAX, 0usize);
|
||||
let mut site = 0usize;
|
||||
for i in &p.instrs {
|
||||
if !i.op.is_load() {
|
||||
continue;
|
||||
}
|
||||
let wsize = ((1u64 << ACCEPT_DATASET_LOG2) >> (i.win as u64).min(2)) as f64;
|
||||
let ratio = d[site] as f64 / (n - n * n / (2.0 * wsize));
|
||||
if ratio < floor {
|
||||
return Err(Reject::LowEntropySite { site: site as u8, distinct: d[site], evaluations: n as u32, ratio_milli: (ratio * 1000.0) as u32 });
|
||||
}
|
||||
if ratio < min.0 {
|
||||
min = (ratio, site);
|
||||
}
|
||||
site += 1;
|
||||
}
|
||||
Ok(min)
|
||||
}
|
||||
|
||||
/// The distinct dataset word indices every load site reads over `units` units of the seed's acceptance stream on
|
||||
/// the closed-form words (the sample caps the count near `units x 32 x 8`, so a site's index entropy is read only
|
||||
/// below about log2 of that).
|
||||
pub fn distinct_indices_v4(p: &Program, units: usize) -> Result<Vec<u32>, Reject> {
|
||||
let loads = p.loads_per_hash();
|
||||
let sites = loads / ITERATIONS;
|
||||
let mut acc = Acc {
|
||||
sources: None,
|
||||
indices: Some(vec![Vec::with_capacity(units * LANES * ITERATIONS); sites]),
|
||||
sat_source: vec![0; sites],
|
||||
and_acc: [u32::MAX; 8],
|
||||
or_acc: [0; 8],
|
||||
saturated: 0,
|
||||
bit_ones: [0; 64],
|
||||
distinct_sum: 0,
|
||||
};
|
||||
let mut lane_addrs = vec![0u32; LANES * loads];
|
||||
for (unit, &base) in accept_base_nonces_n(&p.seed, units).iter().enumerate() {
|
||||
run_unit(p, unit, base, &mut acc, &mut lane_addrs)?;
|
||||
}
|
||||
let mut out = Vec::with_capacity(sites);
|
||||
for ix in acc.indices.take().unwrap().iter_mut() {
|
||||
ix.sort_unstable();
|
||||
ix.dedup();
|
||||
out.push(ix.len() as u32);
|
||||
}
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
/// Whether `class` is the class v4 shape (the 256-instruction shadow block over the class v3 base, the pass count and
|
||||
/// the era set aside): the shape the sub-version 2 rules (a') and (c') apply to, on every draw path.
|
||||
pub fn is_class_v4_shape(class: &LoadClass) -> bool {
|
||||
matches!(class.shadow, Some(ShadowClass { instrs: V4_SHADOW_INSTRS, .. }))
|
||||
&& LoadClass { era: None, shadow: None, ..*class } == LoadClass { shadow: None, ..V4_CLASS }
|
||||
}
|
||||
|
||||
/// One pass of the dataflow freshness over the base program then the shadow block (the order of one iteration),
|
||||
/// from `fresh`; `check` reports the first load that reads a register that is not fresh. The rule (AP-F8-1,
|
||||
/// `docs/analysis/ca3-v4-uniform.md`): a load leaves its destination fresh only if its source was (a saturated
|
||||
/// source reads one fixed word); add, sub, xor, mad and shfl if either operand was; rotl and rotr if the operand
|
||||
/// was (a rotate maps all-ones and zero to themselves); or, mul and mulhi never.
|
||||
fn freshness_pass(p: &Program, fresh: &mut [bool; 8], pair_op: &mut [Option<(Op, usize)>; 8], check: bool) -> Result<(), Reject> {
|
||||
for (k, i) in p.instrs.iter().chain(p.shadow.iter()).enumerate() {
|
||||
let (d, a) = (i.dst as usize, i.src as usize);
|
||||
if check && i.op.is_load() && !fresh[a] {
|
||||
return Err(Reject::UnfreshLoadSource { instr: k as u8, reg: i.src });
|
||||
}
|
||||
// the shared-operand idiom (sub-version 3): or-then-xor or or-then-sub on one operand is `d & ~s`, xor-then-or
|
||||
// is `d | s`: lossy, though the second op would inject on its own (F8's p23: `or r6 |= r4; xor r6 ^= r4`)
|
||||
let masked = matches!((pair_op[d], i.op), (Some((Op::Or, s)), Op::Xor) | (Some((Op::Or, s)), Op::Sub) | (Some((Op::Xor, s)), Op::Or) if s == a);
|
||||
fresh[d] = !masked
|
||||
&& match i.op {
|
||||
Op::Load | Op::WLoad | Op::Scratch | Op::Hot => fresh[a],
|
||||
Op::Add | Op::Sub | Op::Xor | Op::Mad | Op::Shfl => fresh[d] || fresh[a],
|
||||
Op::Rotl | Op::Rotr => fresh[d],
|
||||
Op::Or | Op::Mul | Op::MulHi => false,
|
||||
};
|
||||
pair_op[d] = if matches!(i.op, Op::Or | Op::Xor) && !masked { Some((i.op, a)) } else { None };
|
||||
for r in 0..8 {
|
||||
if r != d {
|
||||
if let Some((_, s)) = pair_op[r] {
|
||||
if s == d {
|
||||
pair_op[r] = None;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Part (a'), class v4 sub-version 2: every load's source is fresh by dataflow in the loop's steady state. The draw
|
||||
/// of `candidate_from_words_class` keeps in-pass sources fresh; this closes the iteration boundary (a source last
|
||||
/// written late in the previous iteration or in the shadow block, which the draw's no-eligible fallback can pick:
|
||||
/// F8's p11, an `or` at 63 feeding a load at 1). The state starts all fresh (the init words are a per-lane hash of
|
||||
/// the nonce) and is run to its fixpoint (it only ever falls, so at most 8 passes change it), then one checking pass.
|
||||
pub fn check_fresh_sources_v4(p: &Program) -> Result<(), Reject> {
|
||||
if !is_class_v4_shape(&p.class) {
|
||||
return Ok(());
|
||||
}
|
||||
let mut fresh = [true; 8];
|
||||
let mut pair_op: [Option<(Op, usize)>; 8] = [None; 8];
|
||||
for _ in 0..9 {
|
||||
let before = (fresh, pair_op);
|
||||
freshness_pass(p, &mut fresh, &mut pair_op, false)?;
|
||||
if (fresh, pair_op) == before {
|
||||
break;
|
||||
}
|
||||
}
|
||||
freshness_pass(p, &mut fresh, &mut pair_op, true)
|
||||
}
|
||||
|
||||
/// Parts (a), (b) and, for class v4 sub-version 2, (a').
|
||||
pub fn check_static(p: &Program) -> Result<(), Reject> {
|
||||
if p.instrs.len() != INSTR_COUNT {
|
||||
panic!("acceptance needs a {INSTR_COUNT}-instruction program");
|
||||
}
|
||||
check_stale_loads(&p.instrs)?;
|
||||
check_injecting_writes(&p.instrs)
|
||||
check_injecting_writes(&p.instrs)?;
|
||||
check_fresh_sources_v4(p)
|
||||
}
|
||||
|
||||
/// The 64 base nonces of the dynamic test for seed words `seed`.
|
||||
pub fn accept_base_nonces(seed: &[u32; 8]) -> [u32; ACCEPT_UNITS] {
|
||||
let v = accept_base_nonces_n(seed, ACCEPT_UNITS);
|
||||
let mut out = [0u32; ACCEPT_UNITS];
|
||||
out.copy_from_slice(&v);
|
||||
out
|
||||
}
|
||||
|
||||
/// The first `n` base nonces of the seed's acceptance stream (the (c) units are the first [`ACCEPT_UNITS`]).
|
||||
pub fn accept_base_nonces_n(seed: &[u32; 8], n: usize) -> Vec<u32> {
|
||||
let mut b = Vec::with_capacity(ACCEPT_TAG.len() + 32);
|
||||
b.extend_from_slice(ACCEPT_TAG);
|
||||
for w in seed {
|
||||
b.extend_from_slice(&w.to_le_bytes());
|
||||
}
|
||||
let mut rng = SplitMix64::new(fnv1a64(&b));
|
||||
let mut out = [0u32; ACCEPT_UNITS];
|
||||
for o in out.iter_mut() {
|
||||
*o = (rng.next() as u32) & !31;
|
||||
}
|
||||
out
|
||||
(0..n).map(|_| (rng.next() as u32) & !31).collect()
|
||||
}
|
||||
|
||||
#[inline(always)]
|
||||
|
|
@ -167,6 +344,13 @@ fn mulhi32(a: u32, b: u32) -> u32 {
|
|||
|
||||
/// Accumulators of the dynamic test over the 64 units.
|
||||
struct Acc {
|
||||
/// (c'') (B): every load's source value per site, recorded when present.
|
||||
sources: Option<Vec<Vec<u32>>>,
|
||||
/// (c'') (A): every load's dataset index per site, recorded when present (the distinct-index pass only).
|
||||
indices: Option<Vec<Vec<u32>>>,
|
||||
/// (c'): per load site (the load's index within the iteration), how many of its evaluations read a source value
|
||||
/// of 0 or all ones (class v4 sub-version 2; counted for every class, judged for class v4 only).
|
||||
sat_source: Vec<u32>,
|
||||
and_acc: [u32; 8],
|
||||
or_acc: [u32; 8],
|
||||
saturated: u32,
|
||||
|
|
@ -198,9 +382,15 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
|
|||
let mut scratch = if p.has_scratch() { Some(ScratchModel::new(p.class.scratch_slots_per_lane())) } else { None };
|
||||
let slot_mask = p.class.scratch_slot_mask();
|
||||
let era = p.class.era;
|
||||
// Class v4 sub-version 3 (AP-F8-3, 7 October 2026): the acceptance interpreter runs the latency-shadow block
|
||||
// after instruction 63 of every iteration, `reps` times with the iteration's `sel`, exactly as the hash does
|
||||
// (verify.rs). Until this commit it ran the 64 base instructions only, so every dynamic test (c) judged a class v4
|
||||
// program the chain never hashes. The shadow block holds no load, so its instructions take the same arms.
|
||||
let shadow_reps = p.shadow_reps();
|
||||
for it in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
for (k, ins) in p.instrs.iter().enumerate() {
|
||||
let shadow_pass = (0..shadow_reps).flat_map(|_| p.shadow.iter().enumerate().map(|(k, i)| (INSTR_COUNT + k, i)));
|
||||
for (k, ins) in p.instrs.iter().enumerate().chain(shadow_pass) {
|
||||
let d = ins.dst as usize;
|
||||
let a = ins.src as usize;
|
||||
match ins.op {
|
||||
|
|
@ -289,8 +479,17 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
|
|||
// word (verify::fold_words); width 1 is the lottery hash's xor of one word.
|
||||
let width = ins.width as usize;
|
||||
let align = !(ins.width as u32 - 1);
|
||||
let site = nload % (loads / ITERATIONS);
|
||||
for lane in 0..LANES {
|
||||
idx[lane] = load_index(era.as_ref(), ins, r[a][lane], mask, ACCEPT_DATASET_LOG2) & align;
|
||||
let v = r[a][lane];
|
||||
acc.sat_source[site] += (v == 0 || v == u32::MAX) as u32;
|
||||
if let Some(src) = acc.sources.as_mut() {
|
||||
src[site].push(v);
|
||||
}
|
||||
idx[lane] = load_index(era.as_ref(), ins, v, mask, ACCEPT_DATASET_LOG2) & align;
|
||||
if let Some(ix) = acc.indices.as_mut() {
|
||||
ix[site].push(idx[lane]);
|
||||
}
|
||||
}
|
||||
if idx.iter().all(|&x| x == idx[0]) {
|
||||
return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 });
|
||||
|
|
@ -368,7 +567,18 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
|
|||
/// Part (c).
|
||||
pub fn check_dynamic(p: &Program) -> Result<AcceptReport, Reject> {
|
||||
let loads = p.loads_per_hash();
|
||||
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let v4 = is_class_v4_shape(&p.class);
|
||||
let sites = loads / ITERATIONS;
|
||||
let mut acc = Acc {
|
||||
sources: None,
|
||||
indices: None,
|
||||
sat_source: vec![0; sites],
|
||||
and_acc: [u32::MAX; 8],
|
||||
or_acc: [0; 8],
|
||||
saturated: 0,
|
||||
bit_ones: [0; 64],
|
||||
distinct_sum: 0,
|
||||
};
|
||||
let mut lane_addrs = vec![0u32; LANES * loads];
|
||||
for (unit, &base) in accept_base_nonces(&p.seed).iter().enumerate() {
|
||||
run_unit(p, unit, base, &mut acc, &mut lane_addrs)?;
|
||||
|
|
@ -382,6 +592,17 @@ pub fn check_dynamic(p: &Program) -> Result<AcceptReport, Reject> {
|
|||
if acc.saturated >= MAX_SATURATED {
|
||||
return Err(Reject::Saturated { count: acc.saturated });
|
||||
}
|
||||
// (c'), class v4 sub-version 2 (AP-F8-1, 7 October 2026): a load whose source is saturated reads one fixed word,
|
||||
// whatever delivered the saturation (an or-written value, a rotate of one, a load after a saturated load); the
|
||||
// source rule of the draw removes the writers it can see and this count catches every delivery. Keyed on the
|
||||
// class v4 shape as the draw's rule is, so v2 and v3 verdicts do not move.
|
||||
if v4 {
|
||||
if let Some((site, &count)) = acc.sat_source.iter().enumerate().find(|(_, &c)| c >= MAX_SATURATED) {
|
||||
return Err(Reject::SaturatedSource { site: site as u8, count });
|
||||
}
|
||||
// (c''), the ratio on the candidate that passed everything else (the draw's last and dearest test)
|
||||
check_distinct_indices_v4(p)?;
|
||||
}
|
||||
let half = (ACCEPT_HASHES / 2) as u32;
|
||||
let mut bias_max = 0u32;
|
||||
for (bit, &ones) in acc.bit_ones.iter().enumerate() {
|
||||
|
|
@ -435,7 +656,7 @@ mod tests {
|
|||
let ds = DatasetSource::from_key(p.seed, DatasetMode::ClosedForm, ACCEPT_DATASET_LOG2);
|
||||
let bases = accept_base_nonces(&p.seed);
|
||||
let loads = p.loads_per_hash();
|
||||
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut acc = Acc { sources: None, indices: None, sat_source: vec![0; 64], and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut la = vec![0u32; LANES * loads];
|
||||
let mut ones = [0u32; 64];
|
||||
for (u, &b) in bases.iter().enumerate() {
|
||||
|
|
@ -450,6 +671,49 @@ mod tests {
|
|||
}
|
||||
}
|
||||
|
||||
/// AP-F8-3 (7 October 2026): the acceptance's execution and `verify.rs` agree on a class v4 program WITH its
|
||||
/// shadow block (the output bit counts over the 64 units on the closed-form dataset, the same sel per iteration),
|
||||
/// so the two paths cannot diverge again: until sub-version 3 the acceptance ran the base program only and judged
|
||||
/// a program the chain never hashes. The devnet epoch-0 seed and the six test eras, 8 x 256 x 27 shadow
|
||||
/// instructions per hash each; the same program with its shadow stripped gives other counts.
|
||||
#[test]
|
||||
fn acceptance_executes_the_shadow_block_as_the_verifier_does() {
|
||||
use crate::generator::{generate_era, EraParams, V3_ALLOWED, V4_CLASS};
|
||||
let hx = |h: &str| -> Vec<u8> { (0..h.len()).step_by(2).map(|i| u8::from_str_radix(&h[i..i + 2], 16).unwrap()).collect() };
|
||||
let g = hx("edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07");
|
||||
let mut eras = vec![g.clone()];
|
||||
for n in 0..6 {
|
||||
eras.push(EraParams::test_era_bytes(&format!("igneum-era-test/{n}")).to_vec());
|
||||
}
|
||||
for era in &eras {
|
||||
let p = generate_era("igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07", &g, V4_CLASS, era, &V3_ALLOWED);
|
||||
assert_eq!(p.shadow.len(), 256);
|
||||
assert_eq!(p.shadow_reps(), 27);
|
||||
let ds = DatasetSource::from_key(p.seed, DatasetMode::ClosedForm, ACCEPT_DATASET_LOG2);
|
||||
let bases = accept_base_nonces(&p.seed);
|
||||
let mut acc = Acc { sources: None, indices: None, sat_source: vec![0; 64], and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut la = vec![0u32; LANES * p.loads_per_hash()];
|
||||
let mut ones = [0u32; 64];
|
||||
for (u, &b) in bases.iter().enumerate() {
|
||||
run_unit(&p, u, b, &mut acc, &mut la).unwrap();
|
||||
for h in crate::verify::hash_warp(&p, b, &ds) {
|
||||
for j in 0..64 {
|
||||
ones[j] += ((h >> j) & 1) as u32;
|
||||
}
|
||||
}
|
||||
}
|
||||
assert_eq!(acc.bit_ones, ones, "the acceptance's execution of a class v4 program (shadow block included) matches the verifier's hashes");
|
||||
// and the same program with its shadow stripped hashes differently: the shadow is executed, not skipped
|
||||
let mut bare = p.clone();
|
||||
bare.shadow.clear();
|
||||
let mut acc2 = Acc { sources: None, indices: None, sat_source: vec![0; 64], and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
for (u, &b) in bases.iter().enumerate() {
|
||||
let _ = run_unit(&bare, u, b, &mut acc2, &mut la);
|
||||
}
|
||||
assert_ne!(acc.bit_ones, acc2.bit_ones, "the shadow block changes the acceptance's execution");
|
||||
}
|
||||
}
|
||||
|
||||
/// The instrumented interpreter agrees with `verify.rs` on the closed-form dataset keyed by the seed words.
|
||||
#[test]
|
||||
fn instrumented_interpreter_matches_verify() {
|
||||
|
|
@ -458,7 +722,7 @@ mod tests {
|
|||
let p = candidate(&s, s.as_bytes(), 0);
|
||||
let ds = DatasetSource::from_key(p.seed, DatasetMode::ClosedForm, ACCEPT_DATASET_LOG2);
|
||||
let bases = accept_base_nonces(&p.seed);
|
||||
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut acc = Acc { sources: None, indices: None, sat_source: vec![0; 64], and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut la = vec![0u32; LANES * p.loads_per_hash()];
|
||||
let mut ones = [0u32; 64];
|
||||
let mut any = false;
|
||||
|
|
@ -501,7 +765,7 @@ mod tests {
|
|||
let p = generate_class("igneum-genesis", LoadClass::hot(96, 4));
|
||||
let words = p.hot_words();
|
||||
let loads = p.loads_per_hash();
|
||||
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut acc = Acc { sources: None, indices: None, sat_source: vec![0; 64], and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut la = vec![0u32; LANES * loads];
|
||||
let mut buckets = [0u64; 16];
|
||||
let mut hot_count = 0u64;
|
||||
|
|
|
|||
|
|
@ -54,6 +54,18 @@ pub const LOAD_SLOTS: usize = 16;
|
|||
/// Attempts before an implementation may treat the seed as a consensus fault (spec 01 section 1.4.6). At the
|
||||
/// measured 5.14 percent rejection rate the chance of 32 consecutive rejections is below 2^-136.
|
||||
pub const MAX_ATTEMPTS: u32 = 32;
|
||||
|
||||
/// The attempt cap of class v4 sub-version 2 (AP-F8-2, 7 October 2026): rules (a') and (c') reject about two thirds of
|
||||
/// candidates, so 32 attempts exhaust with probability about (2/3)^32, 2e-6 per epoch seed, one epoch no node could
|
||||
/// draw every few decades at one epoch an hour (seen at chain-shaped seed igneum-f9/331672). At 256 attempts the
|
||||
/// exhaustion probability is (2/3)^256, under 1e-45; the cost of a rejected attempt is one draw and the 64-unit check,
|
||||
/// about 2 ms on one core, so the worst case is half a second. Keyed on the class v4 shape, so v2 and v3 keep 32.
|
||||
pub const MAX_ATTEMPTS_V4: u32 = 256;
|
||||
|
||||
/// The attempt cap of a class: [`MAX_ATTEMPTS_V4`] for the class v4 shape, [`MAX_ATTEMPTS`] otherwise.
|
||||
pub fn max_attempts_for(class: &LoadClass) -> u32 {
|
||||
if crate::accept::is_class_v4_shape(class) { MAX_ATTEMPTS_V4 } else { MAX_ATTEMPTS }
|
||||
}
|
||||
/// Domain tag of the program id.
|
||||
pub const PROGRAM_ID_TAG: &[u8] = b"igneum-program/";
|
||||
|
||||
|
|
@ -1095,9 +1107,11 @@ pub fn program_id(generator: u32, seed: &[u32; 8], attempt: u32) -> u64 {
|
|||
fnv1a64(&b)
|
||||
}
|
||||
|
||||
/// The sub-version of class v4's program stream, in every generator-4 program id and pack (AP-F8-1: 1 = the
|
||||
/// load-source rule of `candidate_from_words_class`; 0 was the stream of 6 October 2026, never stamped).
|
||||
pub const PROGRAM_SUBVERSION_V4: u16 = 1;
|
||||
/// The sub-version of class v4's program stream, in every generator-4 program id and pack (AP-F8-3: 3 = the
|
||||
/// acceptance executing the shadow block as the hash does, so its verdicts judge the program the chain hashes;
|
||||
/// 2 = the dataflow load-source rule and the saturated-source check (c'), never shipped; 1 = the one-writer rule,
|
||||
/// the stream 0.3.20 and 0.3.21 ship as object byte 5; 0 was the stream of 6 October 2026, never stamped).
|
||||
pub const PROGRAM_SUBVERSION_V4: u16 = 3;
|
||||
|
||||
/// Domain tag of the program id of a read-width class (never collides with [`PROGRAM_ID_TAG`]).
|
||||
pub const PROGRAM_ID_TAG_RW: &[u8] = b"igneum-program-rw/";
|
||||
|
|
@ -1237,19 +1251,26 @@ pub fn candidate_from_words_class(
|
|||
is_hot[slot as usize] = true;
|
||||
}
|
||||
// (2) The instructions. `fresh[r]`: r was written by an earlier instruction and no load has read it since.
|
||||
// Class v4's chain draw (AP-F8-1, 7 October 2026, `docs/analysis/ca3-v4-uniform.md`): a load's source is drawn
|
||||
// only from registers whose last writer injects or is a rotate (`entropy_kept[r]`), never from one last written
|
||||
// by `or`, `mul` or `mulhi` (an `or`-written source is all-ones with probability (3/4)^32 per read and made the
|
||||
// 153x item of the finding). Keyed on the era-composed V4_CLASS so the generator-2 ladder packs keep their stream;
|
||||
// v2 and v3 take no part. The draw order and the stream are otherwise the same, draw for draw.
|
||||
// Keyed on the class with the shadow's pass count set aside (the latency ladder draws the same base program at
|
||||
// every rung: `of_load_class` compares the pass count too and answered None at every rung but 27, found by the
|
||||
// fork's ladder test, 7 October 2026 10:06Z), the same comparison the fork's "v4 is v3 plus the shadow" test makes.
|
||||
let source_rule_v4 = class.era.is_some()
|
||||
&& matches!(class.shadow, Some(ShadowClass { instrs: V4_SHADOW_INSTRS, .. }))
|
||||
// Class v4's load-source rule (AP-F8-1, 7 October 2026, `docs/analysis/ca3-v4-uniform.md`; sub-version 2):
|
||||
// a load's source is drawn only from registers that are FRESH by dataflow. Fresh at the program start (the init
|
||||
// words are a per-lane hash of the nonce); after an op, the destination is fresh when: a load's source was fresh
|
||||
// (a saturated source reads one fixed word and leaves a constant); add, sub, xor, mad or shfl had a fresh operand
|
||||
// (dst or src); rotl or rotr rotated a fresh value (a rotate maps all-ones to all-ones); never after or, mul or
|
||||
// mulhi (an or-written value is all-ones with probability (3/4)^32 per read, the 153x item of the finding; mul
|
||||
// zeroes low bits; mulhi is dense near zero). Sub-version 1 looked one writer back and counted every load and
|
||||
// rotate as fresh, which let an or-saturated value through a rotate or a load-after-load chain (F8's p6, p31).
|
||||
// Keyed on the class v4 shape (the 256-instruction shadow block over the class v3 base, the pass count and the
|
||||
// era set aside) on EVERY draw path, era or not, so a census through candidate_class reads the same stream as
|
||||
// the chain; v2, v3 and every other class take no part. The draw order and the stream are otherwise the same.
|
||||
let source_rule_v4 = matches!(class.shadow, Some(ShadowClass { instrs: V4_SHADOW_INSTRS, .. }))
|
||||
&& LoadClass { era: None, shadow: None, ..class } == LoadClass { shadow: None, ..V4_CLASS };
|
||||
let mut fresh = [false; 8];
|
||||
let mut entropy_kept = [false; 8];
|
||||
let mut fresh_value = [true; 8];
|
||||
// the shared-operand idiom (AP-F8-1, sub-version 3): after `or d |= s`, a later `xor d ^= s` or `sub d -= s` with
|
||||
// s unwritten since is `d & ~s`; after `xor d ^= s`, a later `or d |= s` is `d | s`: lossy either way, though
|
||||
// the second op would count as injecting on its own. `pair_op[d]` holds the (op, s) of the last or/xor on d
|
||||
// while neither d nor s has been written since.
|
||||
let mut pair_op: [Option<(Op, usize)>; 8] = [None; 8];
|
||||
let mut instrs = Vec::with_capacity(INSTR_COUNT);
|
||||
for k in 0..INSTR_COUNT {
|
||||
let mut roll = rng.below(75);
|
||||
|
|
@ -1275,7 +1296,7 @@ pub fn candidate_from_words_class(
|
|||
let mut eligible = [0u64; 8];
|
||||
let mut n = 0usize;
|
||||
for r in 0..8u64 {
|
||||
if r != dst && fresh[r as usize] && (!source_rule_v4 || entropy_kept[r as usize]) {
|
||||
if r != dst && fresh[r as usize] && (!source_rule_v4 || fresh_value[r as usize]) {
|
||||
eligible[n] = r;
|
||||
n += 1;
|
||||
}
|
||||
|
|
@ -1323,7 +1344,26 @@ pub fn candidate_from_words_class(
|
|||
fresh[src as usize] = false;
|
||||
}
|
||||
fresh[dst as usize] = true;
|
||||
entropy_kept[dst as usize] = op.injects() || matches!(op, Op::Rotl | Op::Rotr);
|
||||
let (d, a) = (dst as usize, src as usize);
|
||||
let masked = matches!((pair_op[d], op), (Some((Op::Or, s)), Op::Xor) | (Some((Op::Or, s)), Op::Sub) | (Some((Op::Xor, s)), Op::Or) if s == a);
|
||||
fresh_value[d] = !masked
|
||||
&& match op {
|
||||
Op::Load | Op::WLoad | Op::Scratch | Op::Hot => fresh_value[a],
|
||||
Op::Add | Op::Sub | Op::Xor | Op::Mad | Op::Shfl => fresh_value[d] || fresh_value[a],
|
||||
Op::Rotl | Op::Rotr => fresh_value[d],
|
||||
Op::Or | Op::Mul | Op::MulHi => false,
|
||||
};
|
||||
// a write to d sets or clears d's pair; a write to any register clears every pair that names it as operand
|
||||
pair_op[d] = if matches!(op, Op::Or | Op::Xor) && !masked { Some((op, a)) } else { None };
|
||||
for r in 0..8 {
|
||||
if r != d {
|
||||
if let Some((_, s)) = pair_op[r] {
|
||||
if s == d {
|
||||
pair_op[r] = None;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width, win, off });
|
||||
}
|
||||
// (3) The latency-shadow block (Counter ASIC 3.0 item 8): drawn after the base program from the same stream, so
|
||||
|
|
@ -1410,14 +1450,38 @@ pub fn try_generate_from_seed_bytes(seed_string: &str, seed_bytes: &[u8]) -> Res
|
|||
/// [`try_generate_from_seed_bytes`] for a load class.
|
||||
pub fn try_generate_class(seed_string: &str, seed_bytes: &[u8], class: LoadClass) -> Result<Program, Exhausted> {
|
||||
let mut last = None;
|
||||
for attempt in 0..MAX_ATTEMPTS {
|
||||
let cap = max_attempts_for(&class);
|
||||
for attempt in 0..cap {
|
||||
let p = candidate_class(seed_string, seed_bytes, attempt, class);
|
||||
match check(&p) {
|
||||
Ok(_) => return Ok(p),
|
||||
Err(r) => last = Some(r),
|
||||
}
|
||||
}
|
||||
Err(Exhausted { seed_string: seed_string.to_string(), attempts: MAX_ATTEMPTS, last: last.unwrap() })
|
||||
if crate::accept::is_class_v4_shape(&class) {
|
||||
// AP-F8-2 (7 October 2026, main's ruling: the draw is total and no consensus path panics): a class v4 seed that
|
||||
// exhausts its attempts takes the last-resort program, deterministic and accepted as drawn
|
||||
return Ok(last_resort_v4(candidate_class(seed_string, seed_bytes, cap, class)));
|
||||
}
|
||||
Err(Exhausted { seed_string: seed_string.to_string(), attempts: cap, last: last.unwrap() })
|
||||
}
|
||||
|
||||
/// The last-resort program of a class v4 seed whose [`MAX_ATTEMPTS_V4`] candidates were all rejected (AP-F8-2):
|
||||
/// the candidate at attempt [`MAX_ATTEMPTS_V4`] with every `or`, `mul` and `mulhi` of its base program and its shadow
|
||||
/// block rewritten to `xor` (dst, src and the other fields kept). With no lossy op left every register stays fresh by dataflow from the
|
||||
/// init words on, so rule (a') holds by construction; the program is the seed's consensus program as drawn, with no
|
||||
/// further check, so the draw is total. It is reached with probability about (2/3)^256 per epoch seed (the measured
|
||||
/// (a') plus (c') rejection rate of about two thirds per attempt), under 1e-45: the chain never sees it, and a test
|
||||
/// walks it on real rejected candidates so the path is known to run.
|
||||
pub fn last_resort_v4(mut p: Program) -> Program {
|
||||
// the shadow block runs at the end of every iteration and its own lossy ops feed the next iteration's loads
|
||||
// (rule (a') walks base then shadow to its fixpoint), so both are rewritten
|
||||
for i in p.instrs.iter_mut().chain(p.shadow.iter_mut()) {
|
||||
if matches!(i.op, Op::Or | Op::Mul | Op::MulHi) {
|
||||
i.op = Op::Xor;
|
||||
}
|
||||
}
|
||||
p
|
||||
}
|
||||
|
||||
/// [`try_generate_from_seed_bytes`], treating exhaustion as the consensus fault it is.
|
||||
|
|
@ -1815,6 +1879,184 @@ mod tests {
|
|||
/// from registers whose last writer keeps entropy, so it is its own stream over class v3's load slots and era
|
||||
/// draw; its generator is 4, its id `program_id(4, seed, attempt)` with the sub-version suffix, its era recorded;
|
||||
/// v2 and v3 are untouched.
|
||||
/// Every load's source fresh by dataflow in the loop's steady state: the crate's own rule (a') of `accept.rs`.
|
||||
fn every_load_source_fresh(p: &Program) -> bool {
|
||||
crate::accept::check_fresh_sources_v4(p).is_ok()
|
||||
}
|
||||
|
||||
/// Class v4 sub-version 3, rule (c''): the distinct-index ratio refuses F8's low-entropy band on the chain's own
|
||||
/// candidates (the first attempt of each strong seed past (a), (b), (c) and (c') fails the 2^20 pass), and the
|
||||
/// shared-operand rule removes p23's value constant at the draw: the chain's attempt 1 (id d65122675f16a1c7) draws
|
||||
/// site 7 from r5 and passes; the same program with that source put back to r6 (`or r6 |= r4; xor r6 ^= r4`
|
||||
/// upstream) is refused by the static rule and, run anyway, by the ratio.
|
||||
#[test]
|
||||
fn class_v4_distinct_ratio_rejects_the_low_entropy_band() {
|
||||
use crate::accept::{check_dynamic, check_static, Reject};
|
||||
use crate::seed::seed_words_from_bytes;
|
||||
let w = |s: String| -> Vec<u8> { seed_words_from_bytes(s.as_bytes()).iter().flat_map(|x| x.to_le_bytes()).collect() };
|
||||
let f8 = |k: u32| (w(format!("igneum-attack-f8/program/{k}")), w(format!("igneum-attack-f8/era/{k}")));
|
||||
let candidate = |epoch: &[u8], era: &[u8], attempt: u32| {
|
||||
let label = format!("igneum-epoch/{}", epoch.iter().map(|b| format!("{b:02x}")).collect::<String>());
|
||||
let mut p = candidate_class(&label, epoch, attempt, LoadClass::era(V4_CLASS, era, &V3_ALLOWED));
|
||||
p.generator = GENERATOR_VERSION_V4;
|
||||
p.era_bytes = Some(era.to_vec());
|
||||
p
|
||||
};
|
||||
for k in [15u32, 18, 19, 56] {
|
||||
let (epoch, era) = f8(k);
|
||||
let mut seen = None;
|
||||
for attempt in 0..MAX_ATTEMPTS_V4 {
|
||||
let p = candidate(&epoch, &era, attempt);
|
||||
let t0 = std::time::Instant::now();
|
||||
match check_static(&p).and_then(|_| check_dynamic(&p).map(|_| ())) {
|
||||
Err(Reject::LowEntropySite { site, distinct, evaluations, ratio_milli }) => {
|
||||
println!("p{k} attempt {attempt} id {:016x}: (c'') site {site} read {distinct} distinct over {evaluations}, ratio {}.{:03}, {:.1} s", p.program_id(), ratio_milli / 1000, ratio_milli % 1000, t0.elapsed().as_secs_f64());
|
||||
seen = Some(attempt);
|
||||
break;
|
||||
}
|
||||
Err(_) => continue,
|
||||
Ok(()) => break,
|
||||
}
|
||||
}
|
||||
assert!(seen.is_some(), "p{k}: the chain reaches a candidate only the ratio refuses");
|
||||
}
|
||||
let (epoch, era) = f8(23);
|
||||
let p = candidate(&epoch, &era, 1);
|
||||
assert_eq!(p.program_id(), 0xd65122675f16a1c7);
|
||||
let site7 = p.instrs.iter().enumerate().filter(|(_, i)| i.op.is_load()).nth(7).map(|(k, _)| k).unwrap();
|
||||
println!("p23 attempt 1: site 7 is instruction {site7}, source r{}", p.instrs[site7].src);
|
||||
assert_eq!(p.instrs[site7].src, 5, "the shared-operand rule moved site 7 off r6");
|
||||
assert!(check_static(&p).is_ok(), "p23 attempt 1 passes the static rule");
|
||||
let t0 = std::time::Instant::now();
|
||||
let v = check_dynamic(&p).map(|_| ());
|
||||
println!("p23 attempt 1 dynamic: {:?} in {:.1} s", v.as_ref().err().map(|x| x.to_string()), t0.elapsed().as_secs_f64());
|
||||
assert!(v.is_ok(), "p23 attempt 1 passes the dynamic rule");
|
||||
let mut q = p.clone();
|
||||
q.instrs[site7].src = 6;
|
||||
assert!(matches!(check_static(&q), Err(Reject::UnfreshLoadSource { .. })), "the value constant's load is refused by the static rule");
|
||||
let v = check_dynamic(&q).map(|_| ());
|
||||
println!("p23 attempt 1 with site 7 from r6: {:?}", v.as_ref().err().map(|x| x.to_string()));
|
||||
assert!(matches!(v, Err(Reject::LowEntropySite { .. })), "and by the ratio when run");
|
||||
}
|
||||
|
||||
/// Diagnostic (AP-F8-1, p23): the distinct word indices per site on the closed-form words at 2^20 and 2^24
|
||||
/// evaluations, for the program F8 measured (attempt 1, id d64dbc675f13be9e): site 7 (instruction 38) reads
|
||||
/// r6 = (mulhi(..) | r4) ^ r4 = r6 & ~r4, an andnot idiom the lineage rule counts as fresh.
|
||||
#[test]
|
||||
#[ignore]
|
||||
fn diag_p23_distinct_indices_per_site() {
|
||||
let hx = |h: &str| -> Vec<u8> { (0..h.len()).step_by(2).map(|i| u8::from_str_radix(&h[i..i + 2], 16).unwrap()).collect() };
|
||||
let e = hx("01aa1485fcb5d59223ca40602086e618277decf440188c5a55898f635f7be34f");
|
||||
let r = hx("00951c99e7ef952fd52611b8de7c07cc705cdec9e129a31a19d696c99909f052");
|
||||
let mut p = candidate_class("igneum-epoch/01aa1485fcb5d59223ca40602086e618277decf440188c5a55898f635f7be34f", &e, 1, LoadClass::era(V4_CLASS, &r, &V3_ALLOWED));
|
||||
p.generator = GENERATOR_VERSION_V4;
|
||||
p.era_bytes = Some(r.clone());
|
||||
assert_eq!(p.program_id(), 0xd64dbc675f13be9e);
|
||||
for units in [4096usize, 65536] {
|
||||
let t0 = std::time::Instant::now();
|
||||
let d = crate::accept::distinct_indices_v4(&p, units).unwrap();
|
||||
println!("p23 d64dbc675f13be9e at {} evaluations per site ({:.1} s): distinct word indices per site {:?}; site 7 = {} (log2 {:.1})", units * 32 * 8, t0.elapsed().as_secs_f64(), d, d[7], (d[7] as f64).log2());
|
||||
}
|
||||
}
|
||||
|
||||
/// The threshold measurement for the second sub-version 3 commit: every F8 program (p1 = the devnet epoch-0 seeds,
|
||||
/// p2 to p64 = the attack-pass harness's label-derived seeds), drawn under this commit's verdicts, each load
|
||||
/// site's distinct word indices over 2^20 evaluations against the window expectation N - N^2 / 2W, the minimum
|
||||
/// ratio per seed. Clean seeds set the threshold; the failing seeds of F8's table must sit below it.
|
||||
#[test]
|
||||
#[ignore]
|
||||
fn diag_f8_64_distinct_index_ratios() {
|
||||
use crate::seed::seed_words_from_bytes;
|
||||
let w = |s: String| -> Vec<u8> { seed_words_from_bytes(s.as_bytes()).iter().flat_map(|x| x.to_le_bytes()).collect() };
|
||||
let hx = |h: &str| -> Vec<u8> { (0..h.len()).step_by(2).map(|i| u8::from_str_radix(&h[i..i + 2], 16).unwrap()).collect() };
|
||||
let n = 1u64 << 20;
|
||||
for k in 1..=64u32 {
|
||||
let (epoch, era) = if k == 1 {
|
||||
let g = hx("edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07");
|
||||
(g.clone(), g)
|
||||
} else {
|
||||
(w(format!("igneum-attack-f8/program/{k}")), w(format!("igneum-attack-f8/era/{k}")))
|
||||
};
|
||||
let label = format!("igneum-epoch/{}", epoch.iter().map(|b| format!("{b:02x}")).collect::<String>());
|
||||
let p = generate_era(&label, &epoch, V4_CLASS, &era, &V3_ALLOWED);
|
||||
let t0 = std::time::Instant::now();
|
||||
let d = crate::accept::distinct_indices_v4(&p, 4096).unwrap();
|
||||
let mut ratios = Vec::new();
|
||||
let mut site = 0usize;
|
||||
for i in &p.instrs {
|
||||
if !i.op.is_load() { continue; }
|
||||
let k_off = (i.win as u64).min(2);
|
||||
let wsize = (1u64 << 28) >> k_off;
|
||||
let expected = n as f64 - (n as f64) * (n as f64) / (2.0 * wsize as f64);
|
||||
ratios.push((d[site] as f64 / expected, site, i.win));
|
||||
site += 1;
|
||||
}
|
||||
let (min_ratio, min_site, min_win) = ratios.iter().cloned().fold((9.0, 0, 0), |a, b| if b.0 < a.0 { b } else { a });
|
||||
println!("RATIO p{k} attempt {} id {:016x}: min {:.4} at site {} (win {}) ; all {} ; {:.1} s", p.attempt, p.program_id(), min_ratio, min_site, min_win, ratios.iter().map(|r| format!("{:.3}", r.0)).collect::<Vec<_>>().join(" "), t0.elapsed().as_secs_f64());
|
||||
}
|
||||
}
|
||||
|
||||
/// The 2^24 reach of the ratio for the weak failing seeds (p34, p4, p8, p10 at 1.22x to 1.50x sit inside the
|
||||
/// clean spread at 2^20), with p23 and five clean seeds as the scale.
|
||||
#[test]
|
||||
#[ignore]
|
||||
fn diag_f8_weak_seeds_at_2e24() {
|
||||
use crate::seed::seed_words_from_bytes;
|
||||
let w = |s: String| -> Vec<u8> { seed_words_from_bytes(s.as_bytes()).iter().flat_map(|x| x.to_le_bytes()).collect() };
|
||||
let n = 1u64 << 24;
|
||||
for k in [34u32, 4, 8, 10, 23, 2, 3, 5, 44, 52] {
|
||||
let (epoch, era) = (w(format!("igneum-attack-f8/program/{k}")), w(format!("igneum-attack-f8/era/{k}")));
|
||||
let label = format!("igneum-epoch/{}", epoch.iter().map(|b| format!("{b:02x}")).collect::<String>());
|
||||
let p = generate_era(&label, &epoch, V4_CLASS, &era, &V3_ALLOWED);
|
||||
let t0 = std::time::Instant::now();
|
||||
let d = crate::accept::distinct_indices_v4(&p, 65536).unwrap();
|
||||
let mut ratios = Vec::new();
|
||||
let mut site = 0usize;
|
||||
for i in &p.instrs {
|
||||
if !i.op.is_load() { continue; }
|
||||
let wsize = (1u64 << 28) >> (i.win as u64).min(2);
|
||||
let expected = n as f64 - (n as f64) * (n as f64) / (2.0 * wsize as f64);
|
||||
ratios.push(d[site] as f64 / expected);
|
||||
site += 1;
|
||||
}
|
||||
let min = ratios.iter().cloned().fold(9.0f64, f64::min);
|
||||
println!("RATIO24 p{k} attempt {} id {:016x}: min {:.4} ; all {} ; {:.1} s", p.attempt, p.program_id(), min, ratios.iter().map(|r| format!("{:.3}", r)).collect::<Vec<_>>().join(" "), t0.elapsed().as_secs_f64());
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn class_v4_draw_is_total_with_the_last_resort() {
|
||||
assert_eq!(max_attempts_for(&V4_CLASS), MAX_ATTEMPTS_V4);
|
||||
assert_eq!(max_attempts_for(&LoadClass::era(V4_CLASS, &[7u8; 32], &V3_ALLOWED)), MAX_ATTEMPTS_V4);
|
||||
assert_eq!(max_attempts_for(&V3_CLASS), MAX_ATTEMPTS);
|
||||
assert_eq!(max_attempts_for(&LoadClass::V2), MAX_ATTEMPTS);
|
||||
let class = LoadClass::era(V4_CLASS, &EraParams::test_era_bytes("igneum-era-test/0"), &V3_ALLOWED);
|
||||
// real candidates that rule (a') rejects (the exhausting shape of seed igneum-f9/331672 at 07a809a7: 32 in a
|
||||
// row), repaired by the last resort: every load's source fresh in the loop's steady state
|
||||
let mut rejected = 0;
|
||||
for i in 0..64u32 {
|
||||
let seed = format!("igneum-ca3-v4-amend/total/{i}");
|
||||
for attempt in 0..4u32 {
|
||||
let c = candidate_class(&seed, seed.as_bytes(), attempt, class);
|
||||
if matches!(crate::accept::check_fresh_sources_v4(&c), Err(crate::accept::Reject::UnfreshLoadSource { .. })) {
|
||||
rejected += 1;
|
||||
let fixed = last_resort_v4(c.clone());
|
||||
assert!(crate::accept::check_fresh_sources_v4(&fixed).is_ok(), "{seed} attempt {attempt}: the last resort is fresh at every load");
|
||||
assert!(fixed.instrs.iter().chain(fixed.shadow.iter()).all(|i| !matches!(i.op, Op::Or | Op::Mul | Op::MulHi)));
|
||||
assert_eq!((fixed.instrs.len(), fixed.shadow.len()), (c.instrs.len(), c.shadow.len()));
|
||||
}
|
||||
}
|
||||
}
|
||||
assert!(rejected > 0, "the sample holds (a')-rejected candidates (about two thirds of attempts do)");
|
||||
// the chain path is total: 64 seeds, every one a program
|
||||
for i in 0..64u32 {
|
||||
let seed = format!("igneum-ca3-v4-amend/total/{i}");
|
||||
let p = try_generate_class(&seed, seed.as_bytes(), class).expect("a class v4 seed always draws");
|
||||
assert!(crate::accept::check_fresh_sources_v4(&p).is_ok());
|
||||
assert!(p.attempt < MAX_ATTEMPTS_V4 || p.instrs.iter().chain(p.shadow.iter()).all(|i| !matches!(i.op, Op::Or | Op::Mul | Op::MulHi)));
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn program_class_v4_is_class_v3_with_the_shadow_block() {
|
||||
assert_eq!(V4_CLASS, V3_CLASS.with_shadow(256, 27));
|
||||
|
|
@ -1834,20 +2076,28 @@ mod tests {
|
|||
// whose last writer injects or rotates, so its base program is its own stream (the 6 October stream, equal
|
||||
// to class v3's draw for draw, is sub-version 0 and never stamped); the load slots, the op draws and the era
|
||||
// draw are still class v3's, and every load site obeys the rule
|
||||
assert_eq!(v4.seed, v3.seed);
|
||||
// sub-version 2 rejects candidates (rules (a') and (c')), so the accepted attempt can differ from class v3's and
|
||||
// the seed words carry the attempt: compare with the class v3 candidate at the SAME attempt
|
||||
let v3c = candidate_class("igneum-genesis", b"igneum-genesis", v4.attempt, LoadClass::era(V3_CLASS, &era, &V3_ALLOWED));
|
||||
assert_eq!(v4.seed, v3c.seed);
|
||||
assert_eq!(
|
||||
v4.instrs.iter().map(|i| i.op.is_load()).collect::<Vec<_>>(),
|
||||
v3.instrs.iter().map(|i| i.op.is_load()).collect::<Vec<_>>(),
|
||||
"the load slots are class v3's"
|
||||
v3c.instrs.iter().map(|i| i.op.is_load()).collect::<Vec<_>>(),
|
||||
"the load slots are class v3's at the same attempt"
|
||||
);
|
||||
let mut kept = [false; 8];
|
||||
for (k, i) in v4.instrs.iter().enumerate() {
|
||||
if i.op.is_load() {
|
||||
assert!(kept[i.src as usize], "load #{k} reads r{} whose last writer does not keep entropy", i.src);
|
||||
}
|
||||
kept[i.dst as usize] = i.op.injects() || matches!(i.op, Op::Rotl | Op::Rotr);
|
||||
assert!(every_load_source_fresh(&v4), "every load of the amended v4 program reads a fresh register");
|
||||
let v3c_as_v4 = Program { class: v4.class, shadow: v4.shadow.clone(), ..v3c.clone() };
|
||||
if !every_load_source_fresh(&v3c_as_v4) {
|
||||
assert_ne!(v4.instrs, v3c.instrs, "the amended v4 base program is not class v3's (a lossy-sourced load was redrawn)");
|
||||
}
|
||||
assert_ne!(v4.instrs, v3.instrs, "the amended v4 base program is not class v3's (a lossy-sourced load was redrawn)");
|
||||
// the same seed without an era draws under the rule too (sub-version 2: the rule is keyed on the class shape on
|
||||
// every draw path), and the v3 program of the seed is the known-failed case
|
||||
let v4_no_era = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", V4_CLASS);
|
||||
assert!(every_load_source_fresh(&v4_no_era));
|
||||
// the known-failed case: the v3 program of the same seed and era, re-labelled with the v4 shape so the rule
|
||||
// applies, carries an unfresh load source (96.6 percent of chain-shaped seeds do, ca3-v4-uniform.md section 3)
|
||||
let v3_as_v4 = Program { class: v4.class, shadow: v4.shadow.clone(), ..v3.clone() };
|
||||
println!("class v3 of igneum-genesis under era [7; 32] (attempt {}): every load source fresh = {}; v4 accepted at attempt {}", v3.attempt, every_load_source_fresh(&v3_as_v4), v4.attempt);
|
||||
assert!(v3.shadow.is_empty() && !v3.has_shadow());
|
||||
assert_eq!(v4.shadow.len(), 256);
|
||||
assert_eq!(v4.shadow_reps(), 27);
|
||||
|
|
@ -2105,10 +2355,16 @@ mod tests {
|
|||
fn shadow_class_leaves_the_base_program_and_class_v3_untouched() {
|
||||
// Counter ASIC 3.0 item 8: the shadow is drawn after the 64 base instructions, so the base program, its
|
||||
// attempt and its acceptance verdict are the class's without the shadow; v2 and v3 draw nothing.
|
||||
// Since class v4 sub-version 2 (AP-F8-1) a 256-instruction block over MX8 is the class v4 shape and draws its
|
||||
// load sources under the dataflow rule on every path, so the "untouched" property is shown on a 64-instruction
|
||||
// block (a measurement class, no rule) and the 256-block is shown to obey the rule instead
|
||||
let base = generate_class("igneum-genesis", LoadClass::MX8);
|
||||
let sh64 = generate_class("igneum-genesis", LoadClass::MX8.with_shadow(64, 13));
|
||||
assert_eq!(sh64.instrs, base.instrs);
|
||||
assert_eq!(sh64.attempt, base.attempt);
|
||||
assert_eq!(sh64.shadow.len(), 64);
|
||||
let sh = generate_class("igneum-genesis", LoadClass::MX8.with_shadow(256, 13));
|
||||
assert_eq!(sh.instrs, base.instrs);
|
||||
assert_eq!(sh.attempt, base.attempt);
|
||||
assert!(crate::accept::check_fresh_sources_v4(&sh).is_ok(), "a 256-block over MX8 draws under the class v4 source rule");
|
||||
assert!(base.shadow.is_empty() && !base.has_shadow() && base.shadow_instrs_per_hash() == 0);
|
||||
assert_eq!(sh.shadow.len(), 256);
|
||||
assert!(sh.has_shadow());
|
||||
|
|
|
|||
|
|
@ -18,7 +18,7 @@
|
|||
//! shadow, draw for draw. The edge test is the dataset's alone (the shadow touches no dataset word) and takes no class.
|
||||
|
||||
use igneum_pow::emit::{export_pack, vectors_json};
|
||||
use igneum_pow::generator::{era_generator_of, generate_era, generate_from_seed_bytes, generate_from_seed_bytes_class, generate_from_seed_bytes_program_class, EraParams, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION_V3, GENERATOR_VERSION_V4, INSTR_COUNT, V3_ALLOWED, V3_CLASS};
|
||||
use igneum_pow::generator::{era_generator_of, generate_era, generate_from_seed_bytes, generate_from_seed_bytes_class, generate_from_seed_bytes_program_class, EraParams, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION_V3, GENERATOR_VERSION_V4, INSTR_COUNT, V3_ALLOWED, V3_CLASS, V4_CLASS, V4_SHADOW_INSTRS};
|
||||
use igneum_pow::memhard::{derive_item, mixer, round_key, Cache, MixParams, Shape};
|
||||
use igneum_pow::seed::{day_key, SplitMix64};
|
||||
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
|
||||
|
|
@ -92,16 +92,24 @@ fn contract(p: &Program, seed: &str, class: LoadClass, era: Option<[u8; 32]>) {
|
|||
assert_eq!((i.width, i.win, i.off), (1, 0, 0), "shadow #{k}: no load fields");
|
||||
}
|
||||
let base = program_of(seed, LoadClass { shadow: None, ..class }, era);
|
||||
if p.generator == GENERATOR_VERSION_V4 {
|
||||
// the amended class v4 (AP-F8-1): the chain draw takes a load's source only from registers whose
|
||||
// last writer injects or rotates, so its base program is its own stream, not class v3's; what holds
|
||||
// is the rule itself, checked here on every load site in draw order
|
||||
let mut kept = [false; 8];
|
||||
let v4_shape = LoadClass { era: None, shadow: None, ..class } == LoadClass { shadow: None, ..V4_CLASS } && sh.instrs == V4_SHADOW_INSTRS;
|
||||
if v4_shape {
|
||||
// the amended class v4 (AP-F8-1, sub-version 2): on every draw path a load's source is a register
|
||||
// fresh by dataflow (a load keeps freshness only from a fresh source; add, sub, xor, mad, shfl from
|
||||
// either operand; rotates from their operand; or, mul, mulhi never), so its base program is its own
|
||||
// stream, not class v3's; what holds is the rule itself, checked here on every load site in draw order
|
||||
let mut fresh = [true; 8];
|
||||
for (k, i) in p.instrs.iter().enumerate() {
|
||||
let (d, a) = (i.dst as usize, i.src as usize);
|
||||
if i.op.is_load() {
|
||||
assert!(kept[i.src as usize], "{seed}: load #{k} reads r{} whose last writer does not keep entropy", i.src);
|
||||
assert!(fresh[a], "{seed}: load #{k} reads r{} which is not fresh by dataflow", i.src);
|
||||
}
|
||||
kept[i.dst as usize] = i.op.injects() || matches!(i.op, Op::Rotl | Op::Rotr);
|
||||
fresh[d] = match i.op {
|
||||
Op::Load | Op::WLoad | Op::Scratch | Op::Hot => fresh[a],
|
||||
Op::Add | Op::Sub | Op::Xor | Op::Mad | Op::Shfl => fresh[d] || fresh[a],
|
||||
Op::Rotl | Op::Rotr => fresh[d],
|
||||
Op::Or | Op::Mul | Op::MulHi => false,
|
||||
};
|
||||
}
|
||||
} else {
|
||||
assert_eq!(p.instrs, base.instrs, "{seed}: the base program is the class's without the shadow");
|
||||
|
|
|
|||
|
|
@ -132,9 +132,12 @@ fn program_ids_differ_between_class_v3_and_class_v4_of_one_seed() {
|
|||
let v3 = load("packs-ca2-mixer/mx8-devnet-epoch0", ProgramClass::V3);
|
||||
let v4 = load("packs-ca3-v4/v4-devnet-epoch0", ProgramClass::V4);
|
||||
assert_eq!(v3.reference.program.program_id(), 0x73bc_bfe8_ccf9_88f1, "the v3 control's id as pinned");
|
||||
// the amended class v4 (AP-F8-1, sub-version 1): the id of 6 October 2026, c120d7963abdcd96, is the must-differ
|
||||
// vector (a binary from before the load-source rule), the pinned id below the must-equal one
|
||||
// the amended class v4 (AP-F8-1, AP-F8-3): sub-version 3 (object byte 7; the acceptance executes the shadow block)
|
||||
// is the pinned id below; c120d7963abdcd96 (the 6 October stream, byte 4), 1a4230699a6b9c60 (sub-version 1, the
|
||||
// one-writer rule, 0.3.20's and 0.3.21's byte 5) and a788661687db4bb3 (sub-version 2, never shipped) must differ
|
||||
assert_ne!(v4.reference.program.program_id(), 0xc120_d796_3abd_cd96, "the pre-amendment v4 id must differ");
|
||||
assert_eq!(v4.reference.program.program_id(), 0x1a42_3069_9a6b_9c60, "the amended v4 candidate's id as pinned");
|
||||
assert_ne!(v4.reference.program.program_id(), 0x1a42_3069_9a6b_9c60, "the sub-version-1 id must differ");
|
||||
assert_ne!(v4.reference.program.program_id(), 0xa788_6616_87db_4bb3, "the sub-version-2 id must differ");
|
||||
assert_eq!(v4.reference.program.program_id(), 0xa785_0016_87d8_688a, "the sub-version-3 v4 id as pinned");
|
||||
assert_ne!(v3.reference.program.program_id(), v4.reference.program.program_id());
|
||||
}
|
||||
|
|
|
|||
|
|
@ -210,19 +210,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r0 = r0 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r2 = r2 ^ ds[r7 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r6 = r6 ^ ds[r3 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
|
|
@ -230,13 +230,13 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r7 = r7 ^ ds[r3 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r2 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r7 = r7 ^ ds[r6 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
|
|
@ -248,7 +248,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r3 = r3 ^ ds[r4 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
|
|
@ -257,8 +257,8 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r5 = r5 ^ ds[r3 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
|
|
|
|||
|
|
@ -70,19 +70,19 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r0 = r0 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r2 = r2 ^ ds[r7 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r6 = r6 ^ ds[r3 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
|
|
@ -90,13 +90,13 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r7 = r7 ^ ds[r3 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r2 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r7 = r7 ^ ds[r6 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
|
|
@ -108,7 +108,7 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r3 = r3 ^ ds[r4 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
|
|
@ -117,8 +117,8 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r5 = r5 ^ ds[r3 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
|
|
|
|||
|
|
@ -210,19 +210,19 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r0 = r0 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r2 = r2 ^ ds[r7 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r6 = r6 ^ ds[r3 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
|
|
@ -230,13 +230,13 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r7 = r7 ^ ds[r3 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r2 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r7 = r7 ^ ds[r6 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
|
|
@ -248,7 +248,7 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r3 = r3 ^ ds[r4 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
|
|
@ -257,8 +257,8 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r5 = r5 ^ ds[r3 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
|
|
@ -572,19 +572,19 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon
|
|||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r0 = r0 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r2 = r2 ^ ds[r7 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r6 = r6 ^ ds[r3 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
|
|
@ -592,13 +592,13 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon
|
|||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r7 = r7 ^ ds[r3 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r2 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r7 = r7 ^ ds[r6 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
|
|
@ -610,7 +610,7 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon
|
|||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r3 = r3 ^ ds[r4 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
|
|
@ -619,8 +619,8 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon
|
|||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r5 = r5 ^ ds[r3 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
|
|
|
|||
|
|
@ -46,19 +46,19 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba
|
|||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r0 = r0 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r2 = r2 ^ ds[r7 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r6 = r6 ^ ds[r3 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
|
|
@ -66,13 +66,13 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba
|
|||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r7 = r7 ^ ds[r3 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r2 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r7 = r7 ^ ds[r6 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
|
|
@ -84,7 +84,7 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba
|
|||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r3 = r3 ^ ds[r4 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
|
|
@ -93,8 +93,8 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba
|
|||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r5 = r5 ^ ds[r3 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
|
|
|
|||