docs/plans/funding.md rule 5: the bounty is escrowed and the benchmark live before daily issuance crosses USD 20,000 a day, with the price-per-IGN table per halving period. docs/plans/epoch-length.md section 11: the detector alert held 6 net windows recommends the next shorter ladder step, per-tier costs at 2,400 and 600, the adversary's share per compile class. Section 12: HBM FPGA parts with sources, reads in flight per watt against the RTX 5090 (0.3x to 0.4x at Shuhai's measured rate, up to 1.9x at the unmeasured bank ceiling), the two FPGA lanes in one table, and the ranking: layer 9 above layer 7. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
46 KiB
Epoch length as an era parameter (Counter ASIC 2.0, layer 9)
5 October 2026 (night), branch ca2-epoch, worker "ca2-epoch". Josh: "what about faster program changes?". Status: Designed, with one Measured section (the Mac compile-ahead, section 6) and cited figures for the PC cards. Reserve-only tonight: nothing here changes the devnet's 3,600-DAA-s epoch, no code moves, no consensus effect. The parameter joins the era-parameter table of spec 01 section 1.13.1 beside the layer 4 and 8 draws of docs/plans/era-layout.md (branch ca2-era).
1. What the layer is
Today the program changes every 3,600 DAA s (spec 01 section 1.12, "a prototype value, to be fixed at gate 2"). This layer makes the length a genesis-reserved parameter, epoch_len, base 3,600, settable by 90% miner signal (spec 05 section 5.7) on a fixed ladder between 600 and 7,200 DAA s, with no fork and no release. The point is a chip that must be rebuilt per program (an FPGA with a hard datapath): the shorter the epoch, the smaller the share of each epoch it can mine. Against a GPU the cost is the compile-ahead cadence, measured in section 6, which is what decides the floor.
| Decision | Choice | Why |
|---|---|---|
| How the length is set | 90% signal only; the era stream consumes one draw for it and ignores the value (section 2.4) | a random length buys nothing against the threat and moves two costs (difficulty settle, VDF duty) at random every era |
| Base | 3,600 DAA s | the devnet's hour, unchanged |
| Ladder | 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 (index 0 to 7) | every step divides the day (86,400) and the era (15,552,000), so day and era boundaries are epoch boundaries at every length |
| Where the schedule anchors | the day, not the era (section 2.2) | a change lands at a day boundary, days after the signal, not 180 days after |
| The seed lead | unchanged: checkpoint at least 1,200 DAA s before the epoch start, T_epoch fixed at 600 s on the reference core (section 3) |
the known-program window stays 600 s at every length; the grinding margin and the checkpoint depth are untouched |
| The floor | 600 DAA s (section 6.3) | the slowest measured compile-ahead (the Metal race, 38 s) is 6.3% of 600 s and inside the 600-s window |
2. The parameter
2.1 Definition
epoch_len(d) is the epoch length in DAA seconds in force on day d (spec 01 section 1.12: day d covers [86,400 d, 86,400 (d + 1))). Epoch (d, e) covers DAA scores [86,400 d + epoch_len(d) e, 86,400 d + epoch_len(d) (e + 1)) for e in 0 .. 86,400 / epoch_len(d). Every ladder step divides 86,400, so the last epoch of a day ends exactly at the day boundary and the first epoch of a day starts at it, as 1.12's day key already assumes ("the program seed of the first epoch of day d").
The identifier of an epoch is its start score s = 86,400 d + epoch_len(d) e, not an index. Everything that today takes the epoch index e (the seed checkpoint rule of spec 04 section 4.3 step 1, seed_source in the header, the hot table key of layer 5, the program id) takes s instead. At the base length s = 3,600 e, so nothing changes for the devnet.
"Which program was this block mined under" stays a function of the header alone: the header's DAA score gives d; epoch_len(d) is a function of the chain state before day d - 2 (section 2.3), which is in the header's own past, exactly as the era parameters E_n are; so s follows, then C(s), then the seed. Two nodes validating the same header derive the same s.
2.2 Why the day and not the era
Anchoring to the era would make a change wait up to 180 days, which defeats the purpose ("miners can shorten it when an FPGA appears"). Anchoring to the day makes the response time the signalling window plus two days. The cost is one more thing the day boundary does; it already resets the day key, the cache and the dataset, so the program change at a day boundary is already paid. The parameter still lives in the era table of 1.13.1 (its base, bound and draw slot are genesis constants there); only its activation clock is the day.
2.3 The signal
Spec 05 section 5.7's mechanism, with its open items (window length, delay, field) answered here for this parameter only:
| Item | Value | Note |
|---|---|---|
| Field | 3 bits of the header version (Kaspa's version-bits field), the ladder index 0 to 7; 0 to 7 all valid, the current index is the default a miner carries |
the proposal-id format of O-5.3 stays open for code upgrades; this is a parameter vote, not a code upgrade |
| Window | 7 days of blue blocks (604,800 at 1 block/s, approximate: counted in blocks, not time) | long enough that a rental burst cannot swing it; short enough to answer an FPGA in a week |
| Threshold | 90% of blue blocks in the window carry the same index, and it differs from the index in force | the 90% rule of 5.7; a split vote changes nothing |
| Activation | the first day boundary at least 2 days after the window closes | 2 days covers the 1,200-s seed lead and every compile-ahead measured in section 6 many times over; a node sees the change coming 2 days ahead |
| Hysteresis | none beyond the 90%: to move again, 90% must carry a new index in a fresh window |
epoch_len(d) is therefore: the base (3,600) at genesis; after each activation, the activated index's length. A node derives it from the blue blocks of the window, which are in the header's past.
2.4 The draw slot
The era stream of era-layout.md section 1.1 draws seven values in a fixed order. This parameter takes draw 8: next() is consumed and the value is not used. Reason to consume it: the stream layout is fixed at genesis; if a draw is ever wanted for this parameter it changes no other parameter's value. Reason not to use it: a drawn length answers no threat (an FPGA fleet is defeated by the minimum, not by the variance), and it would move the difficulty-settle share (section 4) and the VDF duty (section 3.3) at random every 180 days.
2.5 The genesis reserve row (spec 01 section 1.13.1)
| Parameter | Base (era 0) | Draw | Bound |
|---|---|---|---|
Epoch length epoch_len |
3,600 DAA s | draw 8 of the era stream consumed, not used; set by 90% signal (section 2.3) at a day boundary | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 |
Unlock condition: live at genesis at the base; the signal path is the unlock. Nothing moves without 90% of blue blocks over 7 days. No height gate is needed because the base is today's value.
3. The seed path at a shorter epoch
3.1 Two options for the VDF
Spec 04 section 4.3: the seed of epoch s is the 10-minute VDF (T_epoch, 600 s on the reference core) of the checkpoint at least 1,200 DAA s before s. The output is known to a reference core 600 s before the epoch starts (a core half as fast finishes at the boundary). So the usable compile-ahead window is 600 s, at every length.
| Option | Rule | Known-program window | Grinding margin (spec 04 section 4.6, against the 2-s window) | Checkpoint depth | VDF duty per core | Verdict |
|---|---|---|---|---|---|---|
| A | T_epoch and the 1,200-s lead stay genesis constants whatever epoch_len |
600 s at every length (lead minus evaluation) | 300x at every length | 1,200 s plus d at every length: a reorg across it stays a merge-depth-scale event (merge depth 3,600 s, spec 04 section 4.3) | 600 / epoch_len: 17% at 3,600, 33% at 1,800, 100% at 600 |
RECOMMENDED |
| B | T = epoch_len / 6, lead = epoch_len / 3 |
epoch_len / 6: 100 s at 600 |
epoch_len / 6 / 2 s: 50x at 600 |
epoch_len / 3: 200 s at 600, a tenth of the merge depth, so a reorg across the seed checkpoint is an ordinary event |
17% at every length | not recommended: the seed checkpoint gets shallow and the margin falls 6x at the floor |
A at the floor: the checkpoint for epoch s is at s - 1,200, two epochs back; the VDF finishes at s - 600, the start of the previous epoch; so each epoch's program is known for exactly one epoch before it starts. "The seed input for s + 1,200 is known during epoch s" is true but useless to a grinder or a compiler: the input does not give the program before the VDF finishes.
3.2 Grinding (proto-vdf/README.md, grinding table)
The table's model is per epoch: a favourable program is worth 3,600 s (1 + a) / (1 + s a) blocks of revenue for the epoch, against one burned block. The benefit is proportional to the epoch length and the cost is not, so without the delay the gain-to-cost ratio falls in proportion: 6.0 to 15.1 : 1 at 3,600 becomes 1.0 to 2.5 : 1 at 600 (the same model, scaled; not re-run) and 12 to 30 : 1 at 7,200. With the delay the gain is 0 at every length, because the grinder learns nothing inside the 2-s window; option A keeps the 300x margin that makes that true. A shorter epoch weakens the attacker's prize and leaves the defence as it is; a longer one raises the prize and the defence still holds at 300x.
3.3 The VDF duty (a consequence, every tier)
Under A a node evaluates a 600-s VDF once per epoch, so one core is busy 600 / epoch_len of the time: 8% at 7,200, 17% at 3,600, 33% at 1,800, 50% at 1,200, 100% at 600. At the floor every mining node spends one core on the VDF without pause (the M5 Max core at 163,000 squarings/s, spec 04 section 4.2; x86 with AVX-512 IFMA faster, open item there). What that means per tier: a pool user, nothing (the pool evaluates); a home miner on a 4-core box, a quarter of the CPU at the floor and the option of spec 04 section 4.7 B (take (y, pi) from a peer, verify in 4.5 ms) if it has no core to spare; a rig, one core for all its cards. Proof traffic: 516 bytes per epoch, 6x per hour at the floor (3 KB an hour); verification 4.5 ms per epoch. The floor is an emergency setting for the day an FPGA appears; at the base nothing changes.
3.4 Fallback (spec 04 section 4.7)
Unchanged. Option B there (no consensus fallback, proofs from peers) matters more at the floor, because a node slower than 2x the reference core and isolated loses up to a whole epoch instead of a sixth of one. Decision at gate 3 as written.
4. The difficulty window as a constraint on the floor
Spec 02: rule v2's reference lane is the newest REF_WINDOW_V2 = 600 DAA s of the epoch, the short lane is 120 chain blocks, and the lanes are epoch-bounded because the hash rate steps by program (35 to 48 Mhash/s across seeds on the M5 Max, spec 01 section 1.12). The constraint:
| Quantity | Rule | At 600 | At 1,800 | At 3,600 | At 7,200 |
|---|---|---|---|---|---|
REF_WINDOW_V2 |
min(600, epoch_len); the lane never spans a program change |
600 (the whole epoch) | 600 | 600 | 600 |
| Settle after an epoch step, measured | 144 s on the +-30% epoch-step scenario (simulator, docs/bench-log.md 3 October difficulty entry; the live v2 settled a 1.85x in-epoch step within 10 minutes, docs/analysis/difficulty-2026-10-04-oscillation.md) |
24% of the epoch | 8% | 4% | 2% |
| Rate error during the settle | bounded by the step (the program-to-program spread, +-15% on the M5 Max; the sign is random per program) |
At the floor the reference lane is the whole epoch and the v2 cap does nothing; the short lane tracks, as the measurement says it does within 10 minutes. The cost is the settle share: with the measured 144 s a 600-s epoch spends about a quarter of its blocks re-settling after each program change, at a rate error up to the program step. Emission follows the block rate (spec 02: E per block), so emission wobbles by up to the step for 144 s per boundary; the sign is random per program, so the drift averages near zero and the variance rises. The 144 s is the +-30% scenario; the +-15% program step is not measured (owed, section 9: sim/difficulty/sim.py has EPOCH = 3600 as a constant). Under the under-10% rule applied to this cost the number lands at 1,800 s, which is why the ladder has steps between the floor and the base: the signal can stop at 1,800 and the floor stays reserved for the case that needs it.
5. The threat it answers, and the one it does not
5.1 An FPGA fleet with a hard datapath per program
A program is 64 instructions x 8 iterations with 16 loads per iteration from a 1 GiB table (spec 01). A fleet that maps each hour's program to a fixed datapath must synthesise, place and route a bitstream per program and load it, inside the 600-s window of section 3.1. Published compile times:
| Source | Device | Figure |
|---|---|---|
| Xiao et al., "Reducing FPGA Compile Time with Separate Compilation for FPGA Building Blocks" (PRflow), FPT 2019, University of Pennsylvania, https://ic.ese.upenn.edu/pdf/prflow_fpt2019.pdf | ZCU102 (XCZU9EG, about 600 k logic cells, approximate) | monolithic Vivado compile of their benchmarks 42 minutes typical, one case 160 minutes; "hour-long compilation times"; their partitioned flow 12 to 18 minutes |
| Aldec, "Save hours of Place & Route time", https://www.aldec.com/en/company/blog/92--save-hours-of-place-and-route-time-in-seconds | Virtex UltraScale class (4.4 M logic cells) | place and route in hours for large designs; UG904's incremental flow about 3x faster when at least 95% of cells and nets are unchanged |
| AMD UG904, Vivado Implementation User Guide (cited through the Aldec blog and the Vivado documentation) | place and route is the longest stage; incremental reuse needs a 95% similar design, which a fresh random program is not |
The window is 600 s at every ladder step (section 3.1); no row above fits it, so a per-program bitstream can never mine the start of an epoch. What the epoch length sets is the share of each epoch the fleet mines once the bitstream lands: max(0, 1 - (compile - 600) / epoch_len).
| Compile time per program | Share of each epoch mined at 600 | at 1,800 | at 3,600 | at 7,200 |
|---|---|---|---|---|
| 12 min (PRflow partitioned, a small kernel) | 0% | 93% | 97% | 98% |
| 42 min (PRflow monolithic) | 0% | 0% | 47% | 74% |
| 160 min (PRflow worst) | 0% | 0% | 0% | 0% |
| hours (Aldec, large parts) | 0% | 0% | 0% | 0% |
Reading: at the base an FPGA fleet with a fast compile farm mines half to all of each hour; at 7,200 a small kernel is a non-issue for it; at 600 nothing it compiles ever runs. That is the lever the signal gives miners. A compile farm does not help a fleet: the window is wall time per program, not throughput, and the same window applies to every board.
5.2 Partial reconfiguration
Loading a partial bitstream is fast: the ICAP takes 4 bytes per cycle at 100 MHz, 400 MB/s, so a region reloads in milliseconds (Microsoft Research, "Minimizing Partial Reconfiguration Overhead with Fully Streaming DMA Engines and Intelligent ICAP Controller", measured 399.6 MB/s, https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/Minimizing20Partial20Reconfiguration20Overhead20with20Fully20Streaming20DMA20Engines20and20Intelligent20ICAP20Controller.pdf; AMD UG909, Dynamic Function eXchange, https://www.amd.com/content/dam/xilinx/support/documents/sw_manuals/xilinx2021_1/ug909-vivado-partial-reconfiguration.pdf). It does not shorten the compile: the partial bitstream is placed and routed by the same flow first. A smaller region compiles faster than the whole part (the PRflow rows above are that effect), which is why the 12-minute row is in the table; the 600-s window still beats it.
5.3 What the layer does not answer
A programmable chip: a soft processor or a GPU-like overlay on an FPGA, or a custom chip with a programmable core, runs any program with no synthesis, and the epoch length does nothing to it. Those are answered by the other layers (the latency bound, the cache and dataset sizes, the mixer, the hot table, the era draws of layers 4 and 8), and an overlay pays area and clock against hard logic (not quantified here; approximate). The public claim does not change by this layer: it removes one route (per-program bitstreams), it does not lower the chip model's headline row.
6. The measurement: compile-ahead per card, and the floor
What a card does between receiving the next seed and swapping (app/igneum-app/src/engine.rs prepare_worker: export the pack from the node; the worker's prepare: generate and compile the program, fill the cache, build the dataset, self-test, optionally race the variants, then swap at the boundary): the per-epoch part is program generation plus kernel compile, the hot table fill (layer 5, per epoch) and the variant race where it runs. The dataset is per day, excluded here; today's worker rebuilds it at every prepare because the pack couples the program to the day, which is a worker change owed before the ladder goes below the base (section 9).
6.1 Per card, measured or cited
| Card, compiler | Program (generate + compile) | Hot table fill (layer 5, docs/plans/hot-table.md, ca2-cache) |
Variant race | Compile-ahead total, race off | Total, race on | Source |
|---|---|---|---|---|---|---|
| Apple M5 Max, Metal | 82 ms (hot-swap entry, 4 October); 58 ms (serve check); 0 to 444 ms over 8 live boundaries (M11 fleet row); 1,798 ms with the Metal compiler cold (variant-racing entry); tonight: 15.9 / 17.7 / 20.4 ms min / median / max over 10 fresh programs, 79 ms for the devnet pack's two libraries (6.2) | 0.07 to 0.22 ms on the GPU | 34.0 / 34.9 / 37.8 s min / median / max over 8 boundaries (M11); 39.8 s one round in the serve check | 0.5 s (1.8 s cold) | 38 s | docs/bench-log.md: "first hourly program swap on the live devnet" (4 October), "miner performance: variant racing" (4 October), M11 table (5 October); section 6.2 |
| NVIDIA RTX 5090, NVRTC 12.8 | NVRTC 151 to 180 ms (M11; 151 ms at the PC 2 14:20 boundary); prepare total 580 to 1,074 ms including cache 68, dataset 113 and self-test 511 ms, so about 0.5 to 1.0 s without the dataset; 1,285 ms on the nvcc path of 4 October | 0.67 ms per 256 MiB cache fill on the 5090 scaled to S / 256, under 1 ms |
232 to 300 ms to compile 17 variants, 112 s of timing for 3 rounds (M11 race rows, PC 1); one round about 37 s (112 / 3, approximate) | 1.0 s | 38 s | docs/bench-log.md M11 table and race rows (5 October), hot-swap entry (4 October), "PC 2 at the 14:20 boundary" (4 October) |
| AMD RX 9070 XT, OpenCL (gfx1201, Adrenalin 26.9.2) | NOT MEASURED at the current worker: proto-opencl/host.c times clBuildProgram only in the prepare path (buildMs in the prepared line) and no prepared line from this card is in any upload (PC 1 app logs 17:26, 18:10, 19:02 UTC; jobs rdna4-serve-4, rdna4-bench-1, run-readwidth-9070-20261005c print cache, dataset and check only) |
not measured on AMD | none: the OpenCL worker has no race (docs/design/miner-tuning.md: 17 names on NVIDIA, 14 on Metal) |
0.31 s plus the compile (cache 9 ms, self-test 300 ms measured, rdna4-serve-4) |
the same | OWED (section 9) |
| AMD Radeon integrated gfx1036 (PC 2), OpenCL | prepare total 6.9 / 9.4 / 11.7 s with the 1 GiB dataset build on the iGPU inside; the compile is not separated | none | under 11.7 s | the same | M11 table; the compile share OWED | |
| Intel UHD (US laptop), OpenCL | OpenCL build 3.0 to 6.4 s; prepare total 7.3 / 7.9 / 11.7 s | none | 6.4 s | the same | M11 table | |
| AMD Radeon integrated gfx1036 (PC 1), OpenCL, beside WSL build jobs | prepare total 55.3 / 115.8 / 123.8 s (the dataset build on a loaded iGPU); 2 of 10 boundaries compiled inline | none | the outlier: 124 s with the dataset | M11 table |
6.2 The Mac, tonight (Measured)
Apple M5 Max (Darwin 25.6.0, 64 GiB), 5 October 2026 21:18 UTC, load average 11 to 14 (other agents' builds and runs; the measure lock held for the 3-s run, with-lock.sh measure bash scratchpad/epoch-measure.sh), proto-metal/igneum-bench built from this branch (e752fc7 + this document) with swiftc -O -target arm64-apple-macos11. Ten distinct programs (seed strings igneum-devnet-v4-epoch0, /epoch1 .. /epoch9, version 2 generator, 128 loads per hash), each generated and compiled at run time (makeLibrary from source plus makeComputePipelineState), dataset 2^28 words, one 2^20 batch and one verify warp per program so the run is the compiles:
./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1
| Compile (library + pipeline), ms | Values over the 10 programs |
|---|---|
| each | 18.8, 17.8, 17.6, 16.1, 18.6, 15.9, 17.6, 17.6, 20.4, 18.2 |
| min / median / max | 15.9 / 17.7 / 20.4 |
| library share | 7.7 to 9.8 ms; pipeline 8.2 to 10.6 ms |
| hash rate during the run (GPU time, loaded Mac) | 27.2 to 30.1 MH/s, all 10 verify warps PASS |
The same pack three times through packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256 (two libraries, memhard.metal and program.metal): compile 79 ms, then 1 ms and 1 ms, the system shader cache answering the identical source; cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU for 1 GiB.
Reading: a fresh program costs this card about 18 ms to compile with the Metal compiler service warm, 79 ms for a pack with the dataset kernels, and up to 1.8 s when the compiler is cold (the variant-racing entry's first seed). The fleet's 0 to 444 ms per boundary (M11) sits between those, so the live figure is the app's cold-start states, not the compile itself. None of it is visible against a 600-s window: the Mac's compile-ahead is the race or nothing.
6.3 Shares and the floor
The usable window is 600 s on the chain (section 3.1: the VDF output lands 600 s before the epoch on a reference core) and 600 DAA s on the devnet (the stand-in lead; the miner sends prepare at about 449 DAA before the boundary, confirm = lead / 4, hot-swap entry).
| Card | Compile-ahead (s) | Share of 600-s epoch | 1,800 | 3,600 | 7,200 | Share of the 600-s window | Fits the devnet's 449-DAA prepare point |
|---|---|---|---|---|---|---|---|
| M5 Max, race off | 0.5 (1.8 cold) | 0.1% (0.3%) | 0.03% | 0.01% | 0.01% | 0.1% | yes |
| M5 Max, race on | 38 | 6.3% | 2.1% | 1.1% | 0.5% | 6.3% | yes |
| RTX 5090, race off | 1.0 | 0.2% | 0.06% | 0.03% | 0.01% | 0.2% | yes |
| RTX 5090, race on (one round) | 38 | 6.3% | 2.1% | 1.1% | 0.5% | 6.3% | yes |
| RX 9070 XT | 0.31 + compile (owed) | ||||||
| Intel UHD | 6.4 (11.7 with the dataset) | 1.1% (2.0%) | 0.4% | 0.2% | 0.1% | 1.1% | yes |
| Radeon integrated, PC 2 | under 11.7 with the dataset | 2.0% | 0.7% | 0.3% | 0.2% | 2.0% | yes |
| Radeon integrated, PC 1 under load | 124 with the dataset | 20.7% | 6.9% | 3.4% | 1.7% | 20.7% | yes, 449 s (but it already compiled inline twice at the base) |
The floor by the rule (the slowest card's compile-ahead under 10% of the epoch and inside the window), dataset excluded: 600 DAA s. The slowest measured compile-ahead is the variant race at about 38 s on both the Mac and the 5090, 6.3% of a 600-s epoch and inside the window. Two conditions carry it:
- The race, if it stays on, costs 6.3% of mining time at the floor against 1.1% at the base. The M11 race rows already found that base wins on both the 5090 and the Mac with the GPU to itself, so the race should default off (or run once a day), which takes the slowest row to 1.0 s. With the race off the floor is set by the Intel iGPU's 6.4-s OpenCL build at 1.1% of 600 s, with the 9070 XT owed.
- The iGPU tier under load (PC 1's 124 s) fails the 10% rule below 1,240 s because it rebuilds the dataset per prepare. The per-day dataset reuse (section 9) removes that; until it ships, that tier compiles inline at the floor and loses the first minutes of each epoch, as it already did twice at the base.
7. Consequences per tier (the standing rule)
| Tier | At the base (3,600) | At the floor (600) | What to do |
|---|---|---|---|
| Home miner, one NVIDIA card, 8 to 32 GB, Windows or Linux | 1 s per hour of prepare, unchanged | 1 s per 10 minutes: 0.2% | nothing; memory unchanged (this parameter adds no bytes) |
| Home miner, one Apple card (M-series), macOS | 0.5 s plus the race's 35 s per hour | race off: 0.5 s per 10 minutes; race on: 6.3% of mining | ship the race default off (M11 finding) |
| Home miner, one AMD card (RDNA 4), Windows or Linux | compile not measured | the same, owed | measure prepared on the 9070 XT (section 9) |
| Integrated GPU (AMD or Intel iGPU) | 7 to 12 s per hour, 55 to 124 s under CPU load | 2% of the epoch, or 21% under load with the per-prepare dataset build | per-day dataset reuse in the worker before any signal below the base |
| Rig (several cards, one node; the app) | one pack export per epoch under EXPORT_LOCK, cards prepare in parallel |
6x the exports per hour, serialised, each sub-second | nothing |
Rig on the installer scripts (packaging/hive/h-run.sh step 4, packaging/linux/bin/igneum-miner.sh step 4, branch rig-install) |
the miner runs with --prepare-packs and --exit-on-seed-change: a worker that prepares swaps in place (the hot-swap entry: rebuilds 0); exit 42 is the fallback when a prepare misses, and it restarts the miner and re-exports the pack (the loaded iGPU missed 2 of 10 boundaries at the base, M11) |
a miss costs a restart plus a pack export plus the inline compile per 10 minutes instead of per hour; a worker with no prepare support (the built-from-source launcher path, engine.rs line 2046) restarts at every boundary, six times an hour |
the rig installer (a3e7b2b03222f5cff): prepare-ahead must be the only boundary path on the rig before any signal below the base; exit 42 stays as the safety net but a miss at the floor is a 10%-of-epoch loss, so the miss causes (the per-prepare dataset build on a loaded iGPU, the prepare sent late) are fixed first (section 9) |
| Pool user | the pool compiles | the same | nothing |
| Every node's CPU | one core 17% busy on the VDF | one core 100% busy | section 3.3; peers' proofs for a node with no core to spare |
| The chain | difficulty settles 144 s per hour (4%) | 24% of blocks in settle, rate error up to the program step | the ladder's middle steps (1,800: 8%) before the floor; the +-15% settle measurement owed |
8. Spec text
8.1 Section 1.12, the epoch row and the paragraph after the table (replacing "Epoch e covers DAA scores ...")
| Clock | Length | What changes | Label |
|---|---|---|---|
| Epoch | epoch_len(d) DAA s, base 3,600; the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200; set by 90% miner signal at a day boundary (section 1.13.1, section 5.7) |
The program: new seed words from the VDF of section 4, new kernel | Designed (Counter ASIC 2.0, layer 9, docs/plans/epoch-length.md); 3,600 stays the prototype value on every network until a signal moves it |
Epoch (d, e) covers DAA scores [86,400 d + L e, 86,400 d + L (e + 1)) with L = epoch_len(d) and e in 0 .. 86,400 / L; every ladder step divides 86,400, so day boundaries are epoch boundaries. The epoch is identified by its start score s = 86,400 d + L e. The epoch of a block is the epoch of its own DAA score, and epoch_len(d) is a function of the blue blocks of the signalling window that closed at least 2 days before day d (section 1.13.1), which are in the header's past; so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch s is generate_from_seed_bytes(program_seed_s) as before, with program_seed_s the VDF output of section 4.3 for the checkpoint at least 1,200 DAA s before s. T_epoch and the 1,200-s lead are genesis constants and do not follow epoch_len: the program of every epoch is known 600 s before it starts on the reference core, at every length. At the base s = 3,600 e and nothing here differs from the previous text.
8.2 Section 1.13.1, one row added to the parameter table, and one paragraph
| Parameter | Base (era 0) | Draw | Bound |
|---|---|---|---|
Epoch length epoch_len |
3,600 DAA s | one draw of the era stream consumed and not used (the value is set by signal) | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 |
epoch_len is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value.
8.3 Section 4.3, one sentence after step 3
The lead and T_epoch are fixed whatever the epoch length of section 1.12: at the floor of 600 DAA s the checkpoint is two epochs back and the program is known one full epoch ahead; at the base it is known for the last sixth of the previous epoch.
9. Owed
| Item | What | How |
|---|---|---|
| RX 9070 XT compile time | clBuildProgram on gfx1201 at the current worker |
a prepared line from the card: PC 1 job with igneum-worker-opencl --serve on the 9070 XT across one boundary, or --bench-pack with a build ms print added beside cache dataset check in host.c (a 2-line change) |
| Radeon integrated compile share | the compile inside the 6.9 to 11.7 s totals | the same print |
| Per-day dataset reuse in the workers | a prepare within the same day keeps the dataset and rebuilds the program only | proto-metal/main.swift, proto-cuda/nvrtc/worker.cpp, proto-opencl/host.c: the pair holds a day key; needed before any signal below the base for the iGPU tier |
| The race default | off, or once a day | the M11 finding; proto-metal/main.swift race = "on" today |
| The rig scripts' boundary path | prepare-ahead as the only path on a rig; exit 42 kept as the net, never the routine | packaging/hive/h-run.sh and packaging/linux/bin/igneum-miner.sh (branch rig-install, a3e7b2b03222f5cff): a prepare-miss counter in the status line, and the built-from-source launcher path (no prepare support) retired before any signal below the base |
| Difficulty settle at a +-15% step and at 600-s epochs | the share of each epoch in settle at the floor | sim/difficulty/sim.py with EPOCH as a parameter and the program-step scenario at +-15% |
| Hot table fill on AMD | the per-epoch fill on the 9070 XT | ca2-cache's PC rows |
| The signal's encoding | the 3-bit field in version against Kaspa's use of the field |
spec 05 section 5.8 |
10. Rows for docs/plans/counter-asic-2.md
The layer table:
| 9 | Epoch length as a signalled era parameter: base 3,600 DAA s, ladder 600 to 7,200, set by 90% signal at a day boundary, lead and T_epoch fixed | a per-program bitstream (an FPGA with a hard datapath): at 600 s nothing it compiles ever runs (42 to 160 min per compile, PRflow FPT 2019; hours on large parts, Aldec) | compile-ahead 1 s per epoch on the 5090, 0.5 s on the M5 Max (38 s with the race on); one CPU core 600 / epoch_len busy on the VDF | reserve-only tonight: docs/plans/epoch-length.md |
The level 3 numbers row:
| Epoch length | 3,600 DAA s at launch; miners can signal it down to 600 (an FPGA defence, no fork) | floor 600: the slowest compile-ahead (the variant race, 38 s) is 6.3% of the epoch and inside the 600-s seed window; at 600 a per-program FPGA bitstream mines 0% of each epoch, at 3,600 up to 47% (42-min compile) | Measured (Mac compile, 5 October), cited (5090, FPGA compile times), owed (9070 XT compile) |
11. The detector as the signal's trigger (Counter ASIC 3.0 item 4)
6 October 2026, worker ca3-detector. The share-pattern detector (tools/observer/detector.mjs, tools/observer/README.md section "Detector") watches the chain for a population of miner ids that behaves like one fixed design: a clique of 3 or more ids whose per-program residual rate vectors correlate above r = 0.8 over a window of 6 closed epochs, with at least one design flag on a member (an excess per-program spread over 10%, no blocks in the first tenth of epochs, a non-uniform nonce pattern, a rate above every known card). It writes live_state.detector and live_events kind detector. The chain does nothing with it; this section is what people do.
| Rule | Value | Why |
|---|---|---|
| Trigger | the detector's alert holds for M = 6 net windows (one window per closed epoch; a window without the candidate counts one down) | 6 windows of 6 epochs span 11 epochs: 11 hours at the base, 110 minutes at the floor; one chance clique never holds that long (README: chance 3-cliques about 0.09 per window at 30 ids, and the alert also needs a design flag) |
| Recommendation | the project publishes the evidence (live_state.detector) and recommends that miners signal the next shorter ladder step: from 3,600 to 2,400, from 2,400 to 1,800, and so on down the ladder of section 1 |
one step at a time, because each step's cost to miners is known (section 7) and the signal takes 7 days plus 2 to land (section 2.3), so a second step can follow the next alert |
| What the chain does | nothing: 90% of blue blocks over 7 days must carry the new index, then the first day boundary 2 days later activates it (section 2.3) | no consensus change, no release, no fork; the detector is advice |
| Reversal | if the alert clears for 7 days at the shorter length, the project recommends signalling back up one step | the shorter epoch costs the chain its difficulty settle share (section 4) and the iGPU tier its compile share (section 7), so it is not kept for nothing |
What the first step (3,600 to 2,400) costs each tier, from section 7's measured compile-aheads scaled to a 2,400-s epoch, and what the floor costs:
| Tier | At 2,400 | At 600 (the floor, after three more steps) | What to do before signalling |
|---|---|---|---|
| Home miner, one NVIDIA card, Windows or Linux | 1 s prepare per 40 min: 0.04% | 0.2% | nothing |
| Home miner, one Apple card, macOS | race off 0.5 s: 0.02%; race on 38 s: 1.6% | 0.1%; 6.3% | ship the race default off (section 9) |
| Home miner, one AMD RDNA 4 card | 0.31 s plus the compile, owed | owed | measure prepared on the 9070 XT (section 9) |
| Integrated GPU (AMD or Intel) | 7 to 12 s: 0.3 to 0.5%; 124 s under CPU load: 5.2% | 2%; 21% | per-day dataset reuse in the worker (section 9) |
| Rig (app, several cards) | one export per 40 min, cards prepare in parallel | 6x the exports an hour | nothing |
| Rig on the installer scripts | a prepare miss costs a restart per 40 min instead of per hour | a miss is 10% of an epoch | prepare-ahead as the only path (section 9) |
| Pool user | nothing | nothing | nothing |
| Every node's CPU (the VDF) | one core 25% busy | one core 100% busy | peers' proofs for a node with no core to spare (section 3.3) |
| The chain | difficulty settles 144 s per epoch: 6% | 24% | the +-15% settle measurement (section 9) |
What the same step costs the adversary, from section 5.1's rule max(0, 1 - (compile - 600) / epoch_len):
| Compile time per program | Share of each epoch mined at 3,600 | at 2,400 (first step) | at 1,800 (second step) | at 600 (the floor) |
|---|---|---|---|---|
| 12 min (PRflow partitioned, a small kernel) | 97% | 95% | 93% | 0% |
| 42 min (PRflow monolithic) | 47% | 20% | 0% | 0% |
| 160 min (PRflow worst) | 0% | 0% | 0% | 0% |
Reading: the first step cuts the 42-minute class from half of every epoch to a fifth and the second step removes it; the 12-minute class (a small kernel on a fast flow) survives every step but the floor, so an alert that persists through two steps is the case the floor was reserved for. A soft overlay compiles nothing and is untouched by every step: that lane is section 12.
12. The FPGA lane and the ranking against layer 7 (Counter ASIC 3.0 item 5)
6 October 2026, worker ca3-detector. No hardware was measured here; every FPGA figure is a product figure or a published measurement with its source and date read, and every derived number is labelled. The GPU side is the measured RTX 5090: 17.5 G dependent 4-byte reads per second at 415 ns (docs/benchmarks/repro.md section 2.2, 6 October 2026) at about 326 W (the bench log's 328.6 W peak, approximate).
12.1 The parts
| Part | HBM | Bandwidth | Stacks, pseudo-channels | Board power | Source (read 6 October 2026) |
|---|---|---|---|---|---|
| AMD Alveo U55C (XCU55, Virtex UltraScale+) | 16 GB HBM2 | 460 GB/s | 2 stacks; 32 pseudo-channels ("32 independent pseudo-channels") | 150 W maximum total, 115 W typical (TDP) | https://www.amd.com/en/products/accelerators/alveo/u55c/a-u55c-p00g-pq-g.html (through a search summary, the page itself timed out); the CoreEL data sheet https://www.c2s.gov.in/Technical_Data_sheet/new/Alveo_U55C_data_sheet.pdf ("HBM Memory 16GB", "HBM Bandwidth 460 GB/s", "Power (TDP) 115W", 1,304K LUTs, 9,024 DSP slices) |
| AMD Alveo U280 (XCU280) | 8 GB HBM2 | 460 GB/s | 2 stacks of 4 GB, each 8 channels of 2 pseudo-channels: 32 pseudo-channels, 32 AXI ports at up to 450 MHz | 225 W total electrical card load | DS963 v1.3 (May 2020) through https://www.digikey.com/en/htmldatasheets/production/3778633/0/0/1/a-u280-a32g-dev-g ; the channel layout from Shuhai (Wang, Huang, Alonso, FCCM 2020, https://arxiv.org/pdf/2005.04324, section II.A and III) |
| AMD Versal HBM (VH1582 class) | 32 GB HBM2e | 819 GB/s | 2 stacks (approximate: the series page gives capacity and bandwidth, not the stack count) | not a board; the VHK158 evaluation kit's power is not published as a product figure (approximate: 150 to 250 W for a card around it) | https://www.amd.com/en/products/adaptive-socs-and-fpgas/versal/hbm-series.html through a search summary ("819 GB/s of memory bandwidth and 32 GB of capacity") |
| Intel (Altera) Agilex 7 M-series | up to 32 GB HBM2e | 820 GB/s (410 per stack) | 2 stacks | not published as a board figure | the Agilex 7 M-series memory-bandwidth white paper, https://www.intel.com/content/dam/www/central-libraries/us/en/documents/2022-12/agilex-7-fpgas-m-series-memory-bandwidth-white-paper.pdf , through a search summary; the product page redirected to a 404 on 6 October 2026 |
12.2 Random reads in flight per watt
The metric of the plan: reads in flight = (random reads per second the HBM controller sustains) x (its random-access latency), divided by board watts; the reads per second per watt column is the same quantity without the latency factor and is the one that sets hash rate per watt.
| Figure | Value | Source and label |
|---|---|---|
| HBM2 idle read latency on the U280, from the FPGA fabric | page hit 106.7 ns, page closed 122.2 ns, page miss 137.8 ns (48 / 55 / 62 cycles at 450 MHz) | Shuhai, Table IV (measured) |
| HBM2 random-access throughput as measured on the U280 at the default address mapping (one bank active per channel, RGBCG) | 2.4 GB/s per AXI channel at 32-byte bursts, 4 KB stride, 256 MB working set = 75 M random reads per second per pseudo-channel, 2.4 G per card over 32 | Shuhai, section IV.C and Figure 7 (measured; the paper's point is that this mapping is the wrong one for random access) |
| HBM2 random-access ceiling, bank-bound, with a bank-interleaved mapping | banks per pseudo-channel / row cycle: 8 to 16 banks / 45 ns = 178 to 356 M activates per second per pseudo-channel, 5.7 to 11.4 G per card over 32 | approximate: tRC 45 ns from JEDEC HBM2 as quoted at https://www.overclock.net/threads/the-hbm2-timings-thread.1743510/ and about 48 ns in the MEMSYS 2018 paper the history cites ([L1]); the banks per pseudo-channel from memory; the tFAW and tRRD command limits are not applied, which makes this a ceiling |
| Queue depth from the fabric | 32 AXI ports; outstanding reads per port in the AMD HBM IP (PG276) not read today | approximate: at 64 per port the fabric holds 2,048 reads in flight, which at 137.8 ns is 14.9 G per second, above the bank ceiling, so the banks bind, not the queues |
| HBM2e parts (Versal HBM, Agilex M) | bandwidth 1.8x the U55C's; the row cycle is the same DRAM | the history's [L1] reading: HBM raises bandwidth, not the row cycle; so the random-read ceiling per stack is the HBM2 ceiling within the clock ratio (approximate) |
| Design | Random reads per second | Latency | Reads in flight | Board W | Reads in flight per W | Reads per second per W | Against the 5090 (per W) |
|---|---|---|---|---|---|---|---|
| RTX 5090, measured | 17.5 G | 415 ns | 7,260 | 326 (approximate) | 22.3 | 53.7 M | 1.0x |
| U55C overlay at Shuhai's measured random rate (default mapping) | 2.4 G | 137.8 ns | 330 | 115 to 150 | 2.2 to 2.9 | 16 to 21 M | 0.30x to 0.39x reads per second per W; 0.10x to 0.13x in flight per W |
| U55C overlay at the bank-bound ceiling (bank-interleaved mapping, approximate) | 5.7 to 11.4 G | 137.8 ns | 790 to 1,570 | 115 to 150 | 5.2 to 13.7 | 38 to 99 M | 0.71x to 1.85x reads per second per W; 0.23x to 0.61x in flight per W |
| U280 overlay, the same ceilings | as the U55C (the same HBM subsystem) | 137.8 ns | the same | 225 | 1.5 to 7.0 | 11 to 51 M | 0.20x to 0.94x reads per second per W |
| Versal HBM or Agilex M overlay (HBM2e, 2 stacks) | the HBM2 ceilings times at most the clock ratio 1.8 (approximate) | about the same | approximate 150 to 250 | under 2x at the ceiling, under 1x at Shuhai's measured rate |
Reading, with every caveat in the table: at the only measured FPGA random-read rate in the literature found today (Shuhai's 2.4 G per second for a two-stack HBM2 card) a soft-overlay FPGA sits at a third of the 5090 per watt, in the RX 9070 XT's class (2.4 to 2.5 G per second, docs/bench-log.md); at the bank-bound ceiling that no published design reaches it could sit between 0.7x and 1.9x per watt, which is why the ceiling row is kept and labelled. The overlay's ALU side is not the bound: a hash is 64 instructions x 8 iterations per lane, 70 G instructions per second at the 5090's 137 MH/s, about 175 soft 32-bit ALUs at 400 MHz on a 1.3 M-LUT part (approximate), and the mixer is not on the hash path. What is owed to turn the ceiling row into a number: one HBM FPGA under a bank-interleaved random-read kernel (a rented U55C or U280 hour, the chase kernel of docs/benchmarks/repro.md ported to an AXI master), measured in reads per second and watts; until then the public claim carries the measured row (0.3x to 0.4x per watt) and names the ceiling.
12.3 Compile-ahead: the two FPGA lanes in one table
| Lane | Compile per program | Share of a 600-s epoch it can mine (section 5.1) | Share of a 3,600-s epoch | Rate against the 5090 per watt | What answers it |
|---|---|---|---|---|---|
| Hard datapath (a bitstream per program) | 42 to 160 min monolithic, 12 to 18 min partitioned (PRflow, FPT 2019, section 5.1) | 0% | 47% at 42 min, 97% at 12 min | a hard datapath has no overlay overhead, but it still reads the same HBM: bounded by the same 0.3x to 1.9x per watt, before its 0% to 97% duty | layer 9: the epoch length (section 11) |
| Soft overlay (a processor on the fabric; the program is data) | none: 0% of every epoch | 100% | 100% | 0.3x to 0.4x per watt measured-basis, up to 1.9x at the unproven bank ceiling | the latency bound itself (the reads per watt), the dataset and cache sizes, the mixer; not the epoch length |
12.4 Layer 9 (epoch length) against layer 7 (the mm8 reserve family): the ranking
| Adversary class | What layer 9 costs it | What layer 7 costs it |
|---|---|---|
| Hard-datapath FPGA | its whole duty at the floor (0% of every epoch; 20% at the first step for the 42-min class, section 11) | nothing: it already compiles a datapath per program, and a new family is one more block in that datapath |
| Soft-overlay FPGA with HBM | nothing (no compile) | nearly nothing: FPGA DSP blocks do int8 dot products natively (the DSP58 of Versal and the DSP48E2 of UltraScale+ carry int8 multiply-accumulate; approximate, from memory), so an mm8 unit is cheap on the fabric |
On-die 256 MiB recompute chip (chip-model-v3.md, 0.92x with the factor) |
nothing (it executes the program) | little: int8 matrix blocks are licensable IP at every node (history section 4.3, addition 6) |
| Partial-store chip with a custom memory system (item 1) | nothing | little, as above |
| One-vendor GPU fleet (the 7.5x AMD gap, item 7) | nothing | it moves the gap: dp4a is 1.17x a step on the 5090, 1.06x on the 9070 XT, 1.6x emulated on Apple (docs/bench-log.md, layer 7 row) |
| User tier | What layer 9 costs it (sections 7 and 11) | What layer 7 costs it (the bench log's layer 7 row, int8-matrix-family.md) |
|---|---|---|
| Home miner, one NVIDIA card (8 to 32 GB), Windows or Linux | 0.04% at 2,400, 0.2% at 600 | nothing relative: dp4a is native, 1.17x a step |
| Home miner, one AMD RDNA 4 card | compile owed; 0.31 s plus it | dp4a native at 1.06x a step: falls 10% behind NVIDIA per step, relatively |
| Home miner, one Apple card, macOS | 0.02% at 2,400 with the race off; 1.6% with it on | the step is emulated at 1.6x the cost: the Apple tier loses share to every native card for as long as the family is live |
| Integrated GPU (AMD, Intel) | 0.3 to 0.5% at 2,400; 5.2% under load until per-day dataset reuse | as the vendor's discrete part (native on AMD, emulated on Intel: not measured) |
| Rig | exports per epoch, parallel prepares | as its cards |
| Pool user | nothing | nothing directly; the pool's card mix shifts |
| Every node (the VDF) | one core 25% busy at 2,400, 100% at 600 | nothing |
| The chain | 6% of each epoch in difficulty settle at 2,400, 24% at 600 | nothing |
| Reversibility | by the same 90% signal, up or down, in 9 days | a family unlocked by height stays; by signal it can be voted off, but the vendors' relative rates are what they are while it is on |
Ranking: layer 9 sits above layer 7 in the reserve. Layer 9 removes one adversary class outright (the per-program bitstream, the first adversary of Lyra2REv2 and X16R in the history), costs every GPU tier under 2% at the first step and under 1% with the race off, and is reversible by the same signal that set it. Layer 7 removes no adversary class (every chip and every FPGA has an int8 dot product; the history's "watch ML hardware"), and its cost lands on one honest tier, Apple at 1.6x per emulated op, with a 10% relative shift against AMD. The chip model does not move for either (the on-die recompute chip's cost is the item derivation): layer 9's value is response time against the FPGA lane, layer 7's is a family a 12-op chip lacks, which the overlay and the chip both acquire cheaply. Item 6 (order the reserve by chip-unfriendliness, mm8 last) follows from the same two tables. Decision for Josh: reserve order layer 9 (epoch length, already live at the base) first, the 32-bit datapath families next, mm8 last; and fund one HBM FPGA hour to replace the ceiling row of 12.2 with a measurement before the public testnet's benchmark page claims a number for this lane.