Merge branch 'ca2-epoch' into ca2-coord

# Conflicts:
#	docs/bench-log.md
This commit is contained in:
igneum-labs 2026-10-05 23:04:46 +00:00
commit b9f0158872
2 changed files with 261 additions and 0 deletions

View file

@ -1665,3 +1665,29 @@ Metal, `with-lock.sh run <script>`, scripts `gpu-a.sh` and `gpu-b.sh` in the ses
| 200 fuzz packs | `--batch-log2 9 --warps 2 --batch-base 4294967040 --batches 1` each | 200/200 PASS, 800/800 standalone units (25,600 hashes), 400/400 in the window; 91 s wall for the 200 runs (20:16:02 to 20:17:33 UTC) |
Totals: 228 of 228 PASS where expected, 3 of 3 FAIL where built in. Reading: the one-warp CPU simulation is exact on Metal under the hosts' present tag policy; the analysis names the host contract (zero the arena at allocation and at the tag counter's wrap, tags from 1, groups a multiple of warps) that turns that into a guarantee, and finds layer 3 does not move the named chip (section 3.4 of the write-up: 2.4x at every share under the cap).
## 5 October 2026 (night), epoch length as an era parameter (Counter ASIC 2.0, layer 9): the Mac's compile-ahead per program
Branch `ca2-epoch`, worker "ca2-epoch"; design and the per-card table in `docs/plans/epoch-length.md`. Question (the project lead: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (`tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh`). `proto-metal/igneum-bench` built from this branch with `swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal` under a build slot.
Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/epoch9`; version 2 generator, 128 loads per hash), each generated and compiled at run time (`makeLibrary` from source plus `makeComputePipelineState`), dataset 2^28 words, one 2^20 batch and one verify warp per program:
./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1
| Program | Compile ms (library + pipeline) | Mhash/s (GPU) | Verify |
|---|---|---|---|
| epoch0 | 18.8 (9.3 + 9.5) | 27.2 | PASS |
| epoch1 | 17.8 (8.8 + 9.1) | 27.9 | PASS |
| epoch2 | 17.6 (8.6 + 9.1) | 27.8 | PASS |
| epoch3 | 16.1 (8.0 + 8.2) | 27.7 | PASS |
| epoch4 | 18.6 (9.0 + 9.6) | 27.5 | PASS |
| epoch5 | 15.9 (7.7 + 8.2) | 28.4 | PASS |
| epoch6 | 17.6 (8.6 + 9.1) | 27.8 | PASS |
| epoch7 | 17.6 (8.7 + 9.0) | 30.1 | PASS |
| epoch8 | 20.4 (9.8 + 10.6) | 28.4 | PASS |
| epoch9 | 18.2 (8.7 + 9.5) | 29.3 | PASS |
| min / median / max | 15.9 / 17.7 / 20.4 | | 10 of 10 |
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.

235
docs/plans/epoch-length.md Normal file
View file

@ -0,0 +1,235 @@
# Epoch length as an era parameter (Counter ASIC 2.0, layer 9)
5 October 2026 (night), branch `ca2-epoch`, worker "ca2-epoch". the project lead: "what about faster program changes?". Status: Designed, with one Measured section (the Mac compile-ahead, section 6) and cited figures for the PC cards. Reserve-only tonight: nothing here changes the devnet's 3,600-DAA-s epoch, no code moves, no consensus effect. The parameter joins the era-parameter table of spec 01 section 1.13.1 beside the layer 4 and 8 draws of `docs/plans/era-layout.md` (branch `ca2-era`).
## 1. What the layer is
Today the program changes every 3,600 DAA s (spec 01 section 1.12, "a prototype value, to be fixed at gate 2"). This layer makes the length a genesis-reserved parameter, `epoch_len`, base 3,600, settable by 90% miner signal (spec 05 section 5.7) on a fixed ladder between 600 and 7,200 DAA s, with no fork and no release. The point is a chip that must be rebuilt per program (an FPGA with a hard datapath): the shorter the epoch, the smaller the share of each epoch it can mine. Against a GPU the cost is the compile-ahead cadence, measured in section 6, which is what decides the floor.
| Decision | Choice | Why |
|---|---|---|
| How the length is set | 90% signal only; the era stream consumes one draw for it and ignores the value (section 2.4) | a random length buys nothing against the threat and moves two costs (difficulty settle, VDF duty) at random every era |
| Base | 3,600 DAA s | the devnet's hour, unchanged |
| Ladder | 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 (index 0 to 7) | every step divides the day (86,400) and the era (15,552,000), so day and era boundaries are epoch boundaries at every length |
| Where the schedule anchors | the day, not the era (section 2.2) | a change lands at a day boundary, days after the signal, not 180 days after |
| The seed lead | unchanged: checkpoint at least 1,200 DAA s before the epoch start, `T_epoch` fixed at 600 s on the reference core (section 3) | the known-program window stays 600 s at every length; the grinding margin and the checkpoint depth are untouched |
| The floor | 600 DAA s (section 6.3) | the slowest measured compile-ahead (the Metal race, 38 s) is 6.3% of 600 s and inside the 600-s window |
## 2. The parameter
### 2.1 Definition
`epoch_len(d)` is the epoch length in DAA seconds in force on day `d` (spec 01 section 1.12: day `d` covers `[86,400 d, 86,400 (d + 1))`). Epoch `(d, e)` covers DAA scores `[86,400 d + epoch_len(d) e, 86,400 d + epoch_len(d) (e + 1))` for `e` in `0 .. 86,400 / epoch_len(d)`. Every ladder step divides 86,400, so the last epoch of a day ends exactly at the day boundary and the first epoch of a day starts at it, as 1.12's day key already assumes ("the program seed of the first epoch of day d").
The identifier of an epoch is its start score `s = 86,400 d + epoch_len(d) e`, not an index. Everything that today takes the epoch index `e` (the seed checkpoint rule of spec 04 section 4.3 step 1, `seed_source` in the header, the hot table key of layer 5, the program id) takes `s` instead. At the base length `s = 3,600 e`, so nothing changes for the devnet.
"Which program was this block mined under" stays a function of the header alone: the header's DAA score gives `d`; `epoch_len(d)` is a function of the chain state before day `d - 2` (section 2.3), which is in the header's own past, exactly as the era parameters `E_n` are; so `s` follows, then `C(s)`, then the seed. Two nodes validating the same header derive the same `s`.
### 2.2 Why the day and not the era
Anchoring to the era would make a change wait up to 180 days, which defeats the purpose ("miners can shorten it when an FPGA appears"). Anchoring to the day makes the response time the signalling window plus two days. The cost is one more thing the day boundary does; it already resets the day key, the cache and the dataset, so the program change at a day boundary is already paid. The parameter still lives in the era table of 1.13.1 (its base, bound and draw slot are genesis constants there); only its activation clock is the day.
### 2.3 The signal
Spec 05 section 5.7's mechanism, with its open items (window length, delay, field) answered here for this parameter only:
| Item | Value | Note |
|---|---|---|
| Field | 3 bits of the header `version` (Kaspa's version-bits field), the ladder index 0 to 7; 0 to 7 all valid, the current index is the default a miner carries | the proposal-id format of O-5.3 stays open for code upgrades; this is a parameter vote, not a code upgrade |
| Window | 7 days of blue blocks (604,800 at 1 block/s, approximate: counted in blocks, not time) | long enough that a rental burst cannot swing it; short enough to answer an FPGA in a week |
| Threshold | 90% of blue blocks in the window carry the same index, and it differs from the index in force | the 90% rule of 5.7; a split vote changes nothing |
| Activation | the first day boundary at least 2 days after the window closes | 2 days covers the 1,200-s seed lead and every compile-ahead measured in section 6 many times over; a node sees the change coming 2 days ahead |
| Hysteresis | none beyond the 90%: to move again, 90% must carry a new index in a fresh window | |
`epoch_len(d)` is therefore: the base (3,600) at genesis; after each activation, the activated index's length. A node derives it from the blue blocks of the window, which are in the header's past.
### 2.4 The draw slot
The era stream of `era-layout.md` section 1.1 draws seven values in a fixed order. This parameter takes draw 8: `next()` is consumed and the value is not used. Reason to consume it: the stream layout is fixed at genesis; if a draw is ever wanted for this parameter it changes no other parameter's value. Reason not to use it: a drawn length answers no threat (an FPGA fleet is defeated by the minimum, not by the variance), and it would move the difficulty-settle share (section 4) and the VDF duty (section 3.3) at random every 180 days.
### 2.5 The genesis reserve row (spec 01 section 1.13.1)
| Parameter | Base (era 0) | Draw | Bound |
|---|---|---|---|
| Epoch length `epoch_len` | 3,600 DAA s | draw 8 of the era stream consumed, not used; set by 90% signal (section 2.3) at a day boundary | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 |
Unlock condition: live at genesis at the base; the signal path is the unlock. Nothing moves without 90% of blue blocks over 7 days. No height gate is needed because the base is today's value.
## 3. The seed path at a shorter epoch
### 3.1 Two options for the VDF
Spec 04 section 4.3: the seed of epoch `s` is the 10-minute VDF (`T_epoch`, 600 s on the reference core) of the checkpoint at least 1,200 DAA s before `s`. The output is known to a reference core 600 s before the epoch starts (a core half as fast finishes at the boundary). So the usable compile-ahead window is 600 s, at every length.
| Option | Rule | Known-program window | Grinding margin (spec 04 section 4.6, against the 2-s window) | Checkpoint depth | VDF duty per core | Verdict |
|---|---|---|---|---|---|---|
| A | `T_epoch` and the 1,200-s lead stay genesis constants whatever `epoch_len` | 600 s at every length (lead minus evaluation) | 300x at every length | 1,200 s plus d at every length: a reorg across it stays a merge-depth-scale event (merge depth 3,600 s, spec 04 section 4.3) | `600 / epoch_len`: 17% at 3,600, 33% at 1,800, 100% at 600 | RECOMMENDED |
| B | `T = epoch_len / 6`, lead `= epoch_len / 3` | `epoch_len / 6`: 100 s at 600 | `epoch_len / 6 / 2 s`: 50x at 600 | `epoch_len / 3`: 200 s at 600, a tenth of the merge depth, so a reorg across the seed checkpoint is an ordinary event | 17% at every length | not recommended: the seed checkpoint gets shallow and the margin falls 6x at the floor |
A at the floor: the checkpoint for epoch `s` is at `s - 1,200`, two epochs back; the VDF finishes at `s - 600`, the start of the previous epoch; so each epoch's program is known for exactly one epoch before it starts. "The seed input for `s + 1,200` is known during epoch `s`" is true but useless to a grinder or a compiler: the input does not give the program before the VDF finishes.
### 3.2 Grinding (proto-vdf/README.md, grinding table)
The table's model is per epoch: a favourable program is worth `3,600 s (1 + a) / (1 + s a)` blocks of revenue for the epoch, against one burned block. The benefit is proportional to the epoch length and the cost is not, so without the delay the gain-to-cost ratio falls in proportion: 6.0 to 15.1 : 1 at 3,600 becomes 1.0 to 2.5 : 1 at 600 (the same model, scaled; not re-run) and 12 to 30 : 1 at 7,200. With the delay the gain is 0 at every length, because the grinder learns nothing inside the 2-s window; option A keeps the 300x margin that makes that true. A shorter epoch weakens the attacker's prize and leaves the defence as it is; a longer one raises the prize and the defence still holds at 300x.
### 3.3 The VDF duty (a consequence, every tier)
Under A a node evaluates a 600-s VDF once per epoch, so one core is busy `600 / epoch_len` of the time: 8% at 7,200, 17% at 3,600, 33% at 1,800, 50% at 1,200, 100% at 600. At the floor every mining node spends one core on the VDF without pause (the M5 Max core at 163,000 squarings/s, spec 04 section 4.2; x86 with AVX-512 IFMA faster, open item there). What that means per tier: a pool user, nothing (the pool evaluates); a home miner on a 4-core box, a quarter of the CPU at the floor and the option of spec 04 section 4.7 B (take `(y, pi)` from a peer, verify in 4.5 ms) if it has no core to spare; a rig, one core for all its cards. Proof traffic: 516 bytes per epoch, 6x per hour at the floor (3 KB an hour); verification 4.5 ms per epoch. The floor is an emergency setting for the day an FPGA appears; at the base nothing changes.
### 3.4 Fallback (spec 04 section 4.7)
Unchanged. Option B there (no consensus fallback, proofs from peers) matters more at the floor, because a node slower than 2x the reference core and isolated loses up to a whole epoch instead of a sixth of one. Decision at gate 3 as written.
## 4. The difficulty window as a constraint on the floor
Spec 02: rule v2's reference lane is the newest `REF_WINDOW_V2 = 600` DAA s of the epoch, the short lane is 120 chain blocks, and the lanes are epoch-bounded because the hash rate steps by program (35 to 48 Mhash/s across seeds on the M5 Max, spec 01 section 1.12). The constraint:
| Quantity | Rule | At 600 | At 1,800 | At 3,600 | At 7,200 |
|---|---|---|---|---|---|
| `REF_WINDOW_V2` | `min(600, epoch_len)`; the lane never spans a program change | 600 (the whole epoch) | 600 | 600 | 600 |
| Settle after an epoch step, measured | 144 s on the +-30% epoch-step scenario (simulator, `docs/bench-log.md` 3 October difficulty entry; the live v2 settled a 1.85x in-epoch step within 10 minutes, `docs/analysis/difficulty-2026-10-04-oscillation.md`) | 24% of the epoch | 8% | 4% | 2% |
| Rate error during the settle | bounded by the step (the program-to-program spread, +-15% on the M5 Max; the sign is random per program) | | | | |
At the floor the reference lane is the whole epoch and the v2 cap does nothing; the short lane tracks, as the measurement says it does within 10 minutes. The cost is the settle share: with the measured 144 s a 600-s epoch spends about a quarter of its blocks re-settling after each program change, at a rate error up to the program step. Emission follows the block rate (spec 02: `E` per block), so emission wobbles by up to the step for 144 s per boundary; the sign is random per program, so the drift averages near zero and the variance rises. The 144 s is the +-30% scenario; the +-15% program step is not measured (owed, section 9: `sim/difficulty/sim.py` has `EPOCH = 3600` as a constant). Under the under-10% rule applied to this cost the number lands at 1,800 s, which is why the ladder has steps between the floor and the base: the signal can stop at 1,800 and the floor stays reserved for the case that needs it.
## 5. The threat it answers, and the one it does not
### 5.1 An FPGA fleet with a hard datapath per program
A program is 64 instructions x 8 iterations with 16 loads per iteration from a 1 GiB table (spec 01). A fleet that maps each hour's program to a fixed datapath must synthesise, place and route a bitstream per program and load it, inside the 600-s window of section 3.1. Published compile times:
| Source | Device | Figure |
|---|---|---|
| Xiao et al., "Reducing FPGA Compile Time with Separate Compilation for FPGA Building Blocks" (PRflow), FPT 2019, University of Pennsylvania, https://ic.ese.upenn.edu/pdf/prflow_fpt2019.pdf | ZCU102 (XCZU9EG, about 600 k logic cells, approximate) | monolithic Vivado compile of their benchmarks 42 minutes typical, one case 160 minutes; "hour-long compilation times"; their partitioned flow 12 to 18 minutes |
| Aldec, "Save hours of Place & Route time", https://www.aldec.com/en/company/blog/92--save-hours-of-place-and-route-time-in-seconds | Virtex UltraScale class (4.4 M logic cells) | place and route in hours for large designs; UG904's incremental flow about 3x faster when at least 95% of cells and nets are unchanged |
| AMD UG904, Vivado Implementation User Guide (cited through the Aldec blog and the Vivado documentation) | | place and route is the longest stage; incremental reuse needs a 95% similar design, which a fresh random program is not |
The window is 600 s at every ladder step (section 3.1); no row above fits it, so a per-program bitstream can never mine the start of an epoch. What the epoch length sets is the share of each epoch the fleet mines once the bitstream lands: `max(0, 1 - (compile - 600) / epoch_len)`.
| Compile time per program | Share of each epoch mined at 600 | at 1,800 | at 3,600 | at 7,200 |
|---|---|---|---|---|
| 12 min (PRflow partitioned, a small kernel) | 0% | 93% | 97% | 98% |
| 42 min (PRflow monolithic) | 0% | 0% | 47% | 74% |
| 160 min (PRflow worst) | 0% | 0% | 0% | 0% |
| hours (Aldec, large parts) | 0% | 0% | 0% | 0% |
Reading: at the base an FPGA fleet with a fast compile farm mines half to all of each hour; at 7,200 a small kernel is a non-issue for it; at 600 nothing it compiles ever runs. That is the lever the signal gives miners. A compile farm does not help a fleet: the window is wall time per program, not throughput, and the same window applies to every board.
### 5.2 Partial reconfiguration
Loading a partial bitstream is fast: the ICAP takes 4 bytes per cycle at 100 MHz, 400 MB/s, so a region reloads in milliseconds (Microsoft Research, "Minimizing Partial Reconfiguration Overhead with Fully Streaming DMA Engines and Intelligent ICAP Controller", measured 399.6 MB/s, https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/Minimizing20Partial20Reconfiguration20Overhead20with20Fully20Streaming20DMA20Engines20and20Intelligent20ICAP20Controller.pdf; AMD UG909, Dynamic Function eXchange, https://www.amd.com/content/dam/xilinx/support/documents/sw_manuals/xilinx2021_1/ug909-vivado-partial-reconfiguration.pdf). It does not shorten the compile: the partial bitstream is placed and routed by the same flow first. A smaller region compiles faster than the whole part (the PRflow rows above are that effect), which is why the 12-minute row is in the table; the 600-s window still beats it.
### 5.3 What the layer does not answer
A programmable chip: a soft processor or a GPU-like overlay on an FPGA, or a custom chip with a programmable core, runs any program with no synthesis, and the epoch length does nothing to it. Those are answered by the other layers (the latency bound, the cache and dataset sizes, the mixer, the hot table, the era draws of layers 4 and 8), and an overlay pays area and clock against hard logic (not quantified here; approximate). The public claim does not change by this layer: it removes one route (per-program bitstreams), it does not lower the chip model's headline row.
## 6. The measurement: compile-ahead per card, and the floor
What a card does between receiving the next seed and swapping (`app/igneum-app/src/engine.rs` `prepare_worker`: export the pack from the node; the worker's `prepare`: generate and compile the program, fill the cache, build the dataset, self-test, optionally race the variants, then swap at the boundary): the per-epoch part is program generation plus kernel compile, the hot table fill (layer 5, per epoch) and the variant race where it runs. The dataset is per day, excluded here; today's worker rebuilds it at every prepare because the pack couples the program to the day, which is a worker change owed before the ladder goes below the base (section 9).
### 6.1 Per card, measured or cited
| Card, compiler | Program (generate + compile) | Hot table fill (layer 5, `docs/plans/hot-table.md`, ca2-cache) | Variant race | Compile-ahead total, race off | Total, race on | Source |
|---|---|---|---|---|---|---|
| Apple M5 Max, Metal | 82 ms (hot-swap entry, 4 October); 58 ms (serve check); 0 to 444 ms over 8 live boundaries (M11 fleet row); 1,798 ms with the Metal compiler cold (variant-racing entry); tonight: 15.9 / 17.7 / 20.4 ms min / median / max over 10 fresh programs, 79 ms for the devnet pack's two libraries (6.2) | 0.07 to 0.22 ms on the GPU | 34.0 / 34.9 / 37.8 s min / median / max over 8 boundaries (M11); 39.8 s one round in the serve check | 0.5 s (1.8 s cold) | 38 s | `docs/bench-log.md`: "first hourly program swap on the live devnet" (4 October), "miner performance: variant racing" (4 October), M11 table (5 October); section 6.2 |
| NVIDIA RTX 5090, NVRTC 12.8 | NVRTC 151 to 180 ms (M11; 151 ms at the PC 2 14:20 boundary); prepare total 580 to 1,074 ms including cache 68, dataset 113 and self-test 511 ms, so about 0.5 to 1.0 s without the dataset; 1,285 ms on the nvcc path of 4 October | 0.67 ms per 256 MiB cache fill on the 5090 scaled to `S / 256`, under 1 ms | 232 to 300 ms to compile 17 variants, 112 s of timing for 3 rounds (M11 race rows, PC 1); one round about 37 s (112 / 3, approximate) | 1.0 s | 38 s | `docs/bench-log.md` M11 table and race rows (5 October), hot-swap entry (4 October), "PC 2 at the 14:20 boundary" (4 October) |
| AMD RX 9070 XT, OpenCL (gfx1201, Adrenalin 26.9.2) | NOT MEASURED at the current worker: `proto-opencl/host.c` times `clBuildProgram` only in the `prepare` path (`buildMs` in the `prepared` line) and no `prepared` line from this card is in any upload (PC 1 app logs 17:26, 18:10, 19:02 UTC; jobs `rdna4-serve-4`, `rdna4-bench-1`, `run-readwidth-9070-20261005c` print cache, dataset and check only) | not measured on AMD | none: the OpenCL worker has no race (`docs/design/miner-tuning.md`: 17 names on NVIDIA, 14 on Metal) | 0.31 s plus the compile (cache 9 ms, self-test 300 ms measured, `rdna4-serve-4`) | the same | OWED (section 9) |
| AMD Radeon integrated gfx1036 (PC 2), OpenCL | prepare total 6.9 / 9.4 / 11.7 s with the 1 GiB dataset build on the iGPU inside; the compile is not separated | | none | under 11.7 s | the same | M11 table; the compile share OWED |
| Intel UHD (US laptop), OpenCL | OpenCL build 3.0 to 6.4 s; prepare total 7.3 / 7.9 / 11.7 s | | none | 6.4 s | the same | M11 table |
| AMD Radeon integrated gfx1036 (PC 1), OpenCL, beside WSL build jobs | prepare total 55.3 / 115.8 / 123.8 s (the dataset build on a loaded iGPU); 2 of 10 boundaries compiled inline | | none | the outlier: 124 s with the dataset | | M11 table |
### 6.2 The Mac, tonight (Measured)
Apple M5 Max (Darwin 25.6.0, 64 GiB), 5 October 2026 21:18 UTC, load average 11 to 14 (other agents' builds and runs; the measure lock held for the 3-s run, `with-lock.sh measure bash scratchpad/epoch-measure.sh`), `proto-metal/igneum-bench` built from this branch (e752fc7 + this document) with `swiftc -O -target arm64-apple-macos11`. Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/epoch9`, version 2 generator, 128 loads per hash), each generated and compiled at run time (`makeLibrary` from source plus `makeComputePipelineState`), dataset 2^28 words, one 2^20 batch and one verify warp per program so the run is the compiles:
./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1
| Compile (library + pipeline), ms | Values over the 10 programs |
|---|---|
| each | 18.8, 17.8, 17.6, 16.1, 18.6, 15.9, 17.6, 17.6, 20.4, 18.2 |
| min / median / max | 15.9 / 17.7 / 20.4 |
| library share | 7.7 to 9.8 ms; pipeline 8.2 to 10.6 ms |
| hash rate during the run (GPU time, loaded Mac) | 27.2 to 30.1 MH/s, all 10 verify warps PASS |
The same pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (two libraries, `memhard.metal` and `program.metal`): compile 79 ms, then 1 ms and 1 ms, the system shader cache answering the identical source; cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU for 1 GiB.
Reading: a fresh program costs this card about 18 ms to compile with the Metal compiler service warm, 79 ms for a pack with the dataset kernels, and up to 1.8 s when the compiler is cold (the variant-racing entry's first seed). The fleet's 0 to 444 ms per boundary (M11) sits between those, so the live figure is the app's cold-start states, not the compile itself. None of it is visible against a 600-s window: the Mac's compile-ahead is the race or nothing.
### 6.3 Shares and the floor
The usable window is 600 s on the chain (section 3.1: the VDF output lands 600 s before the epoch on a reference core) and 600 DAA s on the devnet (the stand-in lead; the miner sends `prepare` at about 449 DAA before the boundary, confirm = lead / 4, hot-swap entry).
| Card | Compile-ahead (s) | Share of 600-s epoch | 1,800 | 3,600 | 7,200 | Share of the 600-s window | Fits the devnet's 449-DAA prepare point |
|---|---|---|---|---|---|---|---|
| M5 Max, race off | 0.5 (1.8 cold) | 0.1% (0.3%) | 0.03% | 0.01% | 0.01% | 0.1% | yes |
| M5 Max, race on | 38 | 6.3% | 2.1% | 1.1% | 0.5% | 6.3% | yes |
| RTX 5090, race off | 1.0 | 0.2% | 0.06% | 0.03% | 0.01% | 0.2% | yes |
| RTX 5090, race on (one round) | 38 | 6.3% | 2.1% | 1.1% | 0.5% | 6.3% | yes |
| RX 9070 XT | 0.31 + compile (owed) | | | | | | |
| Intel UHD | 6.4 (11.7 with the dataset) | 1.1% (2.0%) | 0.4% | 0.2% | 0.1% | 1.1% | yes |
| Radeon integrated, PC 2 | under 11.7 with the dataset | 2.0% | 0.7% | 0.3% | 0.2% | 2.0% | yes |
| Radeon integrated, PC 1 under load | 124 with the dataset | 20.7% | 6.9% | 3.4% | 1.7% | 20.7% | yes, 449 s (but it already compiled inline twice at the base) |
The floor by the rule (the slowest card's compile-ahead under 10% of the epoch and inside the window), dataset excluded: 600 DAA s. The slowest measured compile-ahead is the variant race at about 38 s on both the Mac and the 5090, 6.3% of a 600-s epoch and inside the window. Two conditions carry it:
1. The race, if it stays on, costs 6.3% of mining time at the floor against 1.1% at the base. The M11 race rows already found that base wins on both the 5090 and the Mac with the GPU to itself, so the race should default off (or run once a day), which takes the slowest row to 1.0 s. With the race off the floor is set by the Intel iGPU's 6.4-s OpenCL build at 1.1% of 600 s, with the 9070 XT owed.
2. The iGPU tier under load (PC 1's 124 s) fails the 10% rule below 1,240 s because it rebuilds the dataset per prepare. The per-day dataset reuse (section 9) removes that; until it ships, that tier compiles inline at the floor and loses the first minutes of each epoch, as it already did twice at the base.
## 7. Consequences per tier (the standing rule)
| Tier | At the base (3,600) | At the floor (600) | What to do |
|---|---|---|---|
| Home miner, one NVIDIA card, 8 to 32 GB, Windows or Linux | 1 s per hour of prepare, unchanged | 1 s per 10 minutes: 0.2% | nothing; memory unchanged (this parameter adds no bytes) |
| Home miner, one Apple card (M-series), macOS | 0.5 s plus the race's 35 s per hour | race off: 0.5 s per 10 minutes; race on: 6.3% of mining | ship the race default off (M11 finding) |
| Home miner, one AMD card (RDNA 4), Windows or Linux | compile not measured | the same, owed | measure `prepared` on the 9070 XT (section 9) |
| Integrated GPU (AMD or Intel iGPU) | 7 to 12 s per hour, 55 to 124 s under CPU load | 2% of the epoch, or 21% under load with the per-prepare dataset build | per-day dataset reuse in the worker before any signal below the base |
| Rig (several cards, one node; the app) | one pack export per epoch under `EXPORT_LOCK`, cards prepare in parallel | 6x the exports per hour, serialised, each sub-second | nothing |
| Rig on the installer scripts (`packaging/hive/h-run.sh` step 4, `packaging/linux/bin/igneum-miner.sh` step 4, branch `rig-install`) | the miner runs with `--prepare-packs` and `--exit-on-seed-change`: a worker that prepares swaps in place (the hot-swap entry: `rebuilds 0`); exit 42 is the fallback when a prepare misses, and it restarts the miner and re-exports the pack (the loaded iGPU missed 2 of 10 boundaries at the base, M11) | a miss costs a restart plus a pack export plus the inline compile per 10 minutes instead of per hour; a worker with no prepare support (the built-from-source launcher path, `engine.rs` line 2046) restarts at every boundary, six times an hour | the rig installer (a3e7b2b03222f5cff): prepare-ahead must be the only boundary path on the rig before any signal below the base; exit 42 stays as the safety net but a miss at the floor is a 10%-of-epoch loss, so the miss causes (the per-prepare dataset build on a loaded iGPU, the prepare sent late) are fixed first (section 9) |
| Pool user | the pool compiles | the same | nothing |
| Every node's CPU | one core 17% busy on the VDF | one core 100% busy | section 3.3; peers' proofs for a node with no core to spare |
| The chain | difficulty settles 144 s per hour (4%) | 24% of blocks in settle, rate error up to the program step | the ladder's middle steps (1,800: 8%) before the floor; the +-15% settle measurement owed |
## 8. Spec text
### 8.1 Section 1.12, the epoch row and the paragraph after the table (replacing "Epoch `e` covers DAA scores ...")
| Clock | Length | What changes | Label |
|---|---|---|---|
| Epoch | `epoch_len(d)` DAA s, base 3,600; the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200; set by 90% miner signal at a day boundary (section 1.13.1, section 5.7) | The program: new seed words from the VDF of section 4, new kernel | Designed (Counter ASIC 2.0, layer 9, `docs/plans/epoch-length.md`); 3,600 stays the prototype value on every network until a signal moves it |
Epoch `(d, e)` covers DAA scores `[86,400 d + L e, 86,400 d + L (e + 1))` with `L = epoch_len(d)` and `e` in `0 .. 86,400 / L`; every ladder step divides 86,400, so day boundaries are epoch boundaries. The epoch is identified by its start score `s = 86,400 d + L e`. The epoch of a block is the epoch of its own DAA score, and `epoch_len(d)` is a function of the blue blocks of the signalling window that closed at least 2 days before day `d` (section 1.13.1), which are in the header's past; so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch `s` is `generate_from_seed_bytes(program_seed_s)` as before, with `program_seed_s` the VDF output of section 4.3 for the checkpoint at least 1,200 DAA s before `s`. `T_epoch` and the 1,200-s lead are genesis constants and do not follow `epoch_len`: the program of every epoch is known 600 s before it starts on the reference core, at every length. At the base `s = 3,600 e` and nothing here differs from the previous text.
### 8.2 Section 1.13.1, one row added to the parameter table, and one paragraph
| Parameter | Base (era 0) | Draw | Bound |
|---|---|---|---|
| Epoch length `epoch_len` | 3,600 DAA s | one draw of the era stream consumed and not used (the value is set by signal) | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 |
`epoch_len` is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value.
### 8.3 Section 4.3, one sentence after step 3
The lead and `T_epoch` are fixed whatever the epoch length of section 1.12: at the floor of 600 DAA s the checkpoint is two epochs back and the program is known one full epoch ahead; at the base it is known for the last sixth of the previous epoch.
## 9. Owed
| Item | What | How |
|---|---|---|
| RX 9070 XT compile time | `clBuildProgram` on gfx1201 at the current worker | a `prepared` line from the card: PC 1 job with `igneum-worker-opencl --serve` on the 9070 XT across one boundary, or `--bench-pack` with a `build` ms print added beside `cache dataset check` in `host.c` (a 2-line change) |
| Radeon integrated compile share | the compile inside the 6.9 to 11.7 s totals | the same print |
| Per-day dataset reuse in the workers | a prepare within the same day keeps the dataset and rebuilds the program only | `proto-metal/main.swift`, `proto-cuda/nvrtc/worker.cpp`, `proto-opencl/host.c`: the pair holds a day key; needed before any signal below the base for the iGPU tier |
| The race default | off, or once a day | the M11 finding; `proto-metal/main.swift` `race = "on"` today |
| The rig scripts' boundary path | prepare-ahead as the only path on a rig; exit 42 kept as the net, never the routine | `packaging/hive/h-run.sh` and `packaging/linux/bin/igneum-miner.sh` (branch `rig-install`, a3e7b2b03222f5cff): a prepare-miss counter in the status line, and the built-from-source launcher path (no prepare support) retired before any signal below the base |
| Difficulty settle at a +-15% step and at 600-s epochs | the share of each epoch in settle at the floor | `sim/difficulty/sim.py` with `EPOCH` as a parameter and the program-step scenario at +-15% |
| Hot table fill on AMD | the per-epoch fill on the 9070 XT | ca2-cache's PC rows |
| The signal's encoding | the 3-bit field in `version` against Kaspa's use of the field | spec 05 section 5.8 |
## 10. Rows for `docs/plans/counter-asic-2.md`
The layer table:
| 9 | Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and `T_epoch` fixed | a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) | compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core `600 / epoch_len` busy on the VDF | reserve-only tonight: `docs/plans/epoch-length.md` |
The level 3 numbers row:
| Epoch length | 3,600 DAA s at launch; miners can signal it down to 600 (an FPGA defence, no fork) | floor 600: the slowest compile-ahead (the variant race, 38 s) is 6.3% of the epoch and inside the 600-s seed window; at 600 a per-program FPGA bitstream mines 0% of each epoch, at 3,600 up to 47% (42-min compile) | Measured (Mac compile, 5 October), cited (5090, FPGA compile times), owed (9070 XT compile) |