diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md index e63876400..6ee2a5414 100644 --- a/docs/spec/01-lottery-hash.md +++ b/docs/spec/01-lottery-hash.md @@ -441,7 +441,14 @@ The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One S `epoch_len` is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version, encoding Open in section 5.8) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value. The threat it answers is a per-program hard datapath (an FPGA fleet: 42 to 160 minutes per compile on a mid-size part, PRflow, FPT 2019, hours on large parts; at 600 s nothing it compiles ever runs); it does not answer a programmable chip, which the other layers answer. The floor 600 is set by the slowest compile-ahead measured (the Metal variant race, 38 s on the M5 Max, 6.3% of a 600-s epoch and inside the 600-s seed window; `docs/plans/epoch-length.md` section 6). -The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8) are drawn by the same stream; their rows and the draw order are in `docs/plans/era-layout.md` and enter this table with class v3's vectors (pending its six-era measurement, 5 October 2026). +The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8, decided IN on 5 October 2026, delegated: the six-era hash-rate spread is 1.3% on the RTX 5090, 3.2% on the RX 9070 XT and 0.8% on the M5 Max, under the 5% rule; `docs/plans/era-layout.md`) are drawn under program class v3 by a second stream `S` seeded with words 0 and 1 of `seed_words_from_bytes("igneum-era/" || E_n)` (the index is not in the preimage: `E_n` commits to `n` through the VDF input), seven draws in this order whether or not a value is used: + +1. `W = allowed[below(|allowed|)]`: the width in words of every dataset load of the era, from the genesis-fixed set `allowed`; the set is `{1}` (4 bytes, the read-width decision of 5 October 2026), so the draw is consumed and the width pinned. +2. `M = low32(next()) OR 1`: the stride multiplier, odd, so `x -> x * M` is a bijection. +3. `R = 1 + below(31)`: the stride rotation. +4. to 7. `r_i = next()` for `i` in 0..3: the interleave draws. With `b = log2(W)` and `free = 4 - b`, `c = [b, ..., 15]`; for `i` in `0..free`: `j = i + (r_i mod (16 - b - i))`, swap `c[i]` and `c[j]`; the interleave is `pos = [0, ..., b - 1] ++ sort(c[0..free])`, four ascending bit positions below 16. + +The era parameters are `(W, M, R, pos)`. Dataset mapping under class v3: word `w` holds word `j(w)` of item `t(w)`, where bit `i` of `j(w)` is bit `pos[i]` of `w` and `t(w)` is `w` with bits `pos[0..3]` removed; with `pos = [0, 1, 2, 3]` this is `dataset[w] = item(w >> 4)[w AND 15]` byte for byte; an item keeps its value at every dataset size of at least 2^16 words, and the `W` words of one aligned load lie in one item, so the 4,096-item verifier bound of 1.11 holds. Load address under class v3, for a load site with window draws `(k_off, o)` and a dataset of `2^D` words: `k = min(k_off, D - 26)`, `y = rotl(x * M, R)`, `idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK` (uniform on the window, branch-free, three operations before the mask), one text form in Metal, CUDA and OpenCL. The window draws per instruction (layer 8), after the nine draws of 1.4.3: `k_off = below(3)` (the dataset, a half or a quarter) and `o = low32(next()) AND (2^k_off - 1)`, used only on a load slot, so a class v3 program takes 720 draws; the window never goes below 2^26 words (256 MiB, above the largest on-chip cache in the benchmark) nor above the dataset, and sixteen sites with drawn offsets cover the dataset with high probability (a windows-union census over 300 programs: the SRAM mirror a chip would need is the whole dataset in every hour). The acceptance rule of 1.4.6 is unchanged in its tests and mirrors this address at its constant `D = 28`. Devnet stand-in for `E_n` until the VDF of 4.4 is in the node: era 0 the genesis block hash; era `n >= 1` the hash of the last selected-chain block whose DAA score is below `15,552,000 n - 7,200`. What the interleave buys and does not: a chip that hard-wires one layout reads the wrong 15 words with every word once the era draws another; a chip whose address decoder can permute its address lines pays nothing (stated in the plan). The stride is a bijection with no cryptanalysis yet (Open). "Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, `s[0]` in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured.