Spec 01 and 04: layer 9 epoch_len (1.12, 1.13.1, 4.3), reserve family R1 and the emulation rule (1.13.2), the cache growth rule option C and the step mapping (1.13.3), the mixer_mult row

This commit is contained in:
igneum-josh 2026-10-05 20:52:58 +00:00
parent c9941ca8b6
commit 9b1f8495d0
2 changed files with 17 additions and 5 deletions

View file

@ -396,11 +396,11 @@ All times are DAA seconds since genesis (section 0.6). At 1 block per second one
| Clock | Length | What changes | Label |
|---|---|---|---|
| Epoch | 3,600 DAA s | The program: new seed words from the VDF of section 4, new kernel | Designed (design document, "Always evolving, on three clocks"); the epoch length is a prototype value, to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) |
| Epoch | `epoch_len(d)` DAA s, base 3,600; the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200; set by 90% miner signal at a day boundary (sections 1.13.1 and 5.7) | The program: new seed words from the VDF of section 4, new kernel | Designed (Counter ASIC 2.0 layer 9, 5 October 2026, `docs/plans/epoch-length.md`); 3,600 stays the value on every network until a signal moves it, and stays the prototype value to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) |
| Day | 86,400 DAA s | The day key, hence the cache and the dataset | Designed |
| Era | 15,552,000 DAA s (180 days) | Era parameters and one instruction-family unlock, section 1.13 | Designed; the length is a prototype value (the design says "every 6 months") |
Epoch `e` covers DAA scores `[3,600 e, 3,600 (e + 1))`. The epoch of a block is the epoch of its own DAA score, so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch `e` is `generate_from_seed_bytes(program_seed_e)`: attempt 0 is drawn from `S_e = seed_words_from_bytes(program_seed_e)`, and a rejected attempt is replaced as 1.4.6 says; `program_seed_e` is the 32-byte VDF output of section 4.3.
Epoch `(d, e)` covers DAA scores `[86,400 d + L e, 86,400 d + L (e + 1))` with `L = epoch_len(d)` and `e` in `0 .. 86,400 / L`; every ladder step divides 86,400, so day boundaries are epoch boundaries, and the epoch is identified by its start score `s = 86,400 d + L e`. At the base `L = 3,600` this is `[3,600 e, 3,600 (e + 1))` and nothing below differs from the earlier text. The epoch of a block is the epoch of its own DAA score, and `epoch_len(d)` is a function of the blue blocks of the signalling window that closed at least 2 days before day `d` (section 1.13.1), which are in the header's past, so "which program was this block mined under" is a function of the header alone once the seed is known. `T_epoch` and the 1,200-s lead of section 4.3 are genesis constants and do not follow `epoch_len`: the program of every epoch is known 600 s before it starts on the reference core at every length. The program for epoch `e` is `generate_from_seed_bytes(program_seed_e)`: attempt 0 is drawn from `S_e = seed_words_from_bytes(program_seed_e)`, and a rejected attempt is replaced as 1.4.6 says; `program_seed_e` is the 32-byte VDF output of section 4.3.
Implementation note (devnet, 3 October 2026, `docs/fork-divergence.md` "Epoch seed"): until the VDF of section 4 is in the node, `program_seed_e` is the hash of the last selected-chain block whose DAA score is below `3,600 e - 600`. The 600-DAA-score lead stands in for section 4.3's 20-minute lead: the program of epoch `e` is knowable about 10 minutes before it starts, every block template reports it (`pow_epoch.next_epoch_seed`), and a GPU worker compiles it in the background and swaps at the boundary with no pause (serve protocol `prepare`, `proto-metal/main.swift`, `proto-cuda/host.cu`, `proto-opencl/host.c`). Measured across boundaries on a short-epoch test network in `docs/bench-log.md` (hot-swap entry). The program schedule is a protocol constant; a miner that cannot compile ahead sees the same seed at the same time as everyone else, only later.
@ -422,12 +422,24 @@ The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One S
| Load count | 16 of 64 | not drawn | fixed, so every era is equally memory-bound |
| Output fold rotations | (7, 14, 21), (9, 18, 27) | each `1 + below(31)` | 1..31 |
| Mixer round count | 8 | not drawn | fixed, so the verify budget holds |
| Mixer applications per round `mixer_mult` | 4 (class v3; 1 under class v2) | not drawn | fixed at genesis (Counter ASIC 2.0, 5 October 2026: the M16 recompute chip's only measured lever; section 1.8.5 carries the form; `docs/plans/mixer-x4.md`) |
| Epoch length `epoch_len` | 3,600 DAA s | one draw of the era stream consumed and not used (the value is set by signal) | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 (layer 9, `docs/plans/epoch-length.md`) |
`epoch_len` is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version, encoding Open in section 5.8) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value. The threat it answers is a per-program hard datapath (an FPGA fleet: 42 to 160 minutes per compile on a mid-size part, PRflow, FPT 2019, hours on large parts; at 600 s nothing it compiles ever runs); it does not answer a programmable chip, which the other layers answer. The floor 600 is set by the slowest compile-ahead measured (the Metal variant race, 38 s on the M5 Max, 6.3% of a 600-s epoch and inside the 600-s seed window; `docs/plans/epoch-length.md` section 6).
The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8) are drawn by the same stream; their rows and the draw order are in `docs/plans/era-layout.md` and enter this table with class v3's vectors (pending its six-era measurement, 5 October 2026).
"Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, `s[0]` in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured.
### 1.13.2 Instruction-family reserve
At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era `n >= 1`, reserve family `n` becomes live with weight `W_new` taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by `src AND 31`; bit-field extract with an immediate offset and width; `andn` (`dst = dst AND NOT src`); byte permute of `dst` by an immediate selector; population count and count-leading-zeros folded into `dst` by add; a three-register select (`dst = bit of src2 ? src : dst`); a second shuffle form (`lane + delta mod 32`). The order and `W_new` are Open. A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7.
At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era `n >= 1`, reserve family `n` becomes live with weight `W_new` taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by `src AND 31`; bit-field extract with an immediate offset and width; `andn` (`dst = dst AND NOT src`); byte permute of `dst` by an immediate selector; population count and count-leading-zeros folded into `dst` by add; a three-register select (`dst = bit of src2 ? src : dst`); a second shuffle form (`lane + delta mod 32`). The order and `W_new` are Open, except the first entry, decided 5 October 2026 (Counter ASIC 2.0 layer 7, delegated; Josh confirms for the public testnet genesis; `docs/analysis/int8-matrix-family.md`):
> Reserve family R1, `mm8` (integer matrix). Semantics: section 2.2 of `docs/analysis/int8-matrix-family.md`, uint8 operands from `src` and `src2` in the m8n8k16 fragment layout, one int32 element of C per lane selected by the immediate `bit`, added into `dst` modulo 2^32. Weight at unlock `W_new = 4` points, taken proportionally from the ten live non-load families (the load weight and count are untouched). Edge vectors, each a hand-built unit run on every vendor: all bytes 0xFF in A and B (C = 1,040,400 everywhere); all bytes 0x80 (C = 262,144); A all zero (C = 0); `dst` = 0xFFFFFFFF with a nonzero C (the wrap); alternating 0x00 and 0xFF by lane; `bit` = 0 and 1 on the same fragments. Unlock: at the start of era n = 4 (DAA 62,208,000), or earlier by the 90% signalling path of section 5.7; never by a release. Native paths: PTX `mma.sync` `.u8` (sm_75+), AMD WMMA `i32_16x16x16_iu8` (RDNA 3 and 4), Metal 4 `mpp::tensor_ops::matmul2d` (`uchar x uchar -> int`); the per-lane `dot4` form is emulation on Apple (1.6x per op unsigned, measured 5 October 2026) and is not the reserved form.
A vendor that can only emulate. A family enters the reserve when it is bit-exact on every vendor of 1.15. A vendor that reaches the result only by emulation (no instruction or library path) does not block entry if the measured penalty of the emulation on that vendor, on the family's own probe (a dependent chain of the op against the same vendor's integer ALU chain), is at most 8x per op, AND the family's weight at unlock keeps the emulating vendor's hash-rate loss under 5% on the memory-hard hash, checked on the vendor's card with the family live. A family whose emulation exceeds either bound stays out of the reserve until the vendor ships a path.
A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7.
### 1.13.3 Dataset growth
@ -437,7 +449,7 @@ Designed: 2 GiB at genesis plus 0.5 GiB per year (design document, "Which cards
N_d = floor((2 GiB + 0.5 GiB * (86,400 d / 31,536,000)) / 64 bytes)
```
evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day and is recomputed with the day key. Two consequences are Open:
evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day on this average and is recomputed with the day key. Decided 5 October 2026 (Counter ASIC 2.0 layer 6, delegated; Josh confirms for the public testnet genesis): the cache grows with the dataset, doubling when the dataset doubles: `cache_log2_words(d) = 26 + growth_doublings(d)`, `growth_doublings(d) = floor(log2(1 + d / 1,460))` for day `d` since genesis (doublings at years 4, 12, 28 and 60), so the cache is 256 MiB at genesis, 512 MiB from year 4, 1 GiB from year 12; the verifier's one-core fill is 0.2, 0.4 and 0.8 s at those steps (0.2 s per 256 MiB, section 1.12), under 1 s at every step of the schedule. The dataset steps to the next power of two on the same doublings (option (b) below, recommended to Josh with the card-lifetime consequences in `docs/analysis/card-lifetime-2026-10-05.md`: a 4 GB card mines to year 4, an 8 GB card to year 12, a 12 GB card to year 28 with the cache freed after the daily build). Why the cache grows at all: an SRAM mirror of a flat 256 MiB cache is about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density (AMD V-Cache, 64 MB on 41 mm^2 at 7 nm; `docs/analysis/sram-mirror.md`), so the cache size never prices a chip out; its job is to stay above any GPU's on-die cache (96 MB on the RTX 5090, 128 MB on GB202), which a flat 256 MiB loses within the decade. The GPU frees the cache after the daily dataset build; the hash never reads it. Two consequences remain Open:
- Index mapping. `src AND MASK` requires a power-of-two size. For a non-power-of-two `N_d` the proposed mapping is `idx = (src * N_words) >> 32` computed in 64 bits (a multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only). At `N_words = 2^28` this gives `src >> 4`, not `src AND MASK`, so adopting it changes the 1 GiB vectors; gate 1 chooses between (a) the multiply-shift mapping with new vectors, or (b) power-of-two sizes only, growing in steps (2 GiB, 4 GiB) on the same schedule's average, which keeps `AND MASK` and means a 4 GiB card lasts until the 4 GiB step instead of fading.
- The item index `t` is 32 bits, so the construction as written tops out at 2^32 items = 256 GiB, which the schedule reaches after 508 years. No action needed.

View file

@ -49,7 +49,7 @@ Designed. Epoch e is the DAA-score interval `[3,600 e, 3,600 (e + 1))` (section
1. **Seed checkpoint.** `C(e)` is the highest-index checkpoint (section 3, C1) whose checkpoint block has DAA score at most `3,600 e - 1,200`: the latest checkpoint at least 20 minutes of DAA time before the epoch starts. The checkpoint block hash is Kaspa's full header hash, which covers the nonce, as the grinding defence requires (`proto-vdf/README.md`: a hash that covers only the body would let a miner start the VDF while still searching nonces).
2. **Evaluation.** `input = hash(C(e))`, `T = T_epoch`, run 4.2. `program_seed_e` is the 32-byte output; `proof_e` is (T, y, pi).
3. **Program.** `S_e = seed_words_from_bytes(program_seed_e)` (section 1.3.1), program = `generate_from_words(S_e)` (section 1.4). The 1,200-s lead is 2x the reference evaluation time, so a core half as fast as the reference still finishes before the epoch (4.6).
3. **Program.** `S_e = seed_words_from_bytes(program_seed_e)` (section 1.3.1), program = `generate_from_words(S_e)` (section 1.4). The 1,200-s lead is 2x the reference evaluation time, so a core half as fast as the reference still finishes before the epoch (4.6). The lead and `T_epoch` are fixed whatever the epoch length of section 1.12: at the floor of 600 DAA s the checkpoint is two epochs back and the program is known one full epoch ahead; at the base it is known for the last sixth of the previous epoch (`docs/plans/epoch-length.md`, section 3).
4. **Header.** Every header carries `seed_source = hash(C(e))` for its own epoch (section 2.4). A header is valid under the lottery only if its `seed_source` is a block on its own selected chain at the blue score of checkpoint index `i(C(e))`, and its PoW verifies under the program derived from that block. The proof is not in the header: a node verifies `proof_e` once per epoch (4.5 ms) and caches `S_e`.
Determinism of step 1 is the point of the rule: which block is "the checkpoint at blue score 30 i" is a function of the header's own past, so two nodes validating the same header derive the same program, and a header mined under a reorged-away checkpoint names a block that is not on its chain and is invalid. Whether `C(e)` must be certified (section 3) or merely be the selected-chain block at that blue score in the header's past is Open (O-4.3): requiring certification couples mining to finality liveness (a stall longer than the lead would stop the program from being derivable), which the design document accepts ("an epoch cannot start without a valid proof") and this specification argues against, because the chain is meant to keep running on plain GHOSTDAG through a finality pause (section 3.7 item 2). The proposal: the selected-chain block at that blue score, certified or not, deep enough (1,200 DAA s plus d) that a reorg across it is a merge-depth-scale event.