diff --git a/docs/plans/epoch-length.md b/docs/plans/epoch-length.md index 149f90994..caf8b52ce 100644 --- a/docs/plans/epoch-length.md +++ b/docs/plans/epoch-length.md @@ -233,3 +233,104 @@ The layer table: The level 3 numbers row: | Epoch length | 3,600 DAA s at launch; miners can signal it down to 600 (an FPGA defence, no fork) | floor 600: the slowest compile-ahead (the variant race, 38 s) is 6.3% of the epoch and inside the 600-s seed window; at 600 a per-program FPGA bitstream mines 0% of each epoch, at 3,600 up to 47% (42-min compile) | Measured (Mac compile, 5 October), cited (5090, FPGA compile times), owed (9070 XT compile) | + +## 11. The detector as the signal's trigger (Counter ASIC 3.0 item 4) + +6 October 2026, worker ca3-detector. The share-pattern detector (`tools/observer/detector.mjs`, `tools/observer/README.md` section "Detector") watches the chain for a population of miner ids that behaves like one fixed design: a clique of 3 or more ids whose per-program residual rate vectors correlate above r = 0.8 over a window of 6 closed epochs, with at least one design flag on a member (an excess per-program spread over 10%, no blocks in the first tenth of epochs, a non-uniform nonce pattern, a rate above every known card). It writes `live_state.detector` and `live_events` kind `detector`. The chain does nothing with it; this section is what people do. + +| Rule | Value | Why | +|---|---|---| +| Trigger | the detector's alert holds for M = 6 net windows (one window per closed epoch; a window without the candidate counts one down) | 6 windows of 6 epochs span 11 epochs: 11 hours at the base, 110 minutes at the floor; one chance clique never holds that long (README: chance 3-cliques about 0.09 per window at 30 ids, and the alert also needs a design flag) | +| Recommendation | the project publishes the evidence (`live_state.detector`) and recommends that miners signal the next shorter ladder step: from 3,600 to 2,400, from 2,400 to 1,800, and so on down the ladder of section 1 | one step at a time, because each step's cost to miners is known (section 7) and the signal takes 7 days plus 2 to land (section 2.3), so a second step can follow the next alert | +| What the chain does | nothing: 90% of blue blocks over 7 days must carry the new index, then the first day boundary 2 days later activates it (section 2.3) | no consensus change, no release, no fork; the detector is advice | +| Reversal | if the alert clears for 7 days at the shorter length, the project recommends signalling back up one step | the shorter epoch costs the chain its difficulty settle share (section 4) and the iGPU tier its compile share (section 7), so it is not kept for nothing | + +What the first step (3,600 to 2,400) costs each tier, from section 7's measured compile-aheads scaled to a 2,400-s epoch, and what the floor costs: + +| Tier | At 2,400 | At 600 (the floor, after three more steps) | What to do before signalling | +|---|---|---|---| +| Home miner, one NVIDIA card, Windows or Linux | 1 s prepare per 40 min: 0.04% | 0.2% | nothing | +| Home miner, one Apple card, macOS | race off 0.5 s: 0.02%; race on 38 s: 1.6% | 0.1%; 6.3% | ship the race default off (section 9) | +| Home miner, one AMD RDNA 4 card | 0.31 s plus the compile, owed | owed | measure `prepared` on the 9070 XT (section 9) | +| Integrated GPU (AMD or Intel) | 7 to 12 s: 0.3 to 0.5%; 124 s under CPU load: 5.2% | 2%; 21% | per-day dataset reuse in the worker (section 9) | +| Rig (app, several cards) | one export per 40 min, cards prepare in parallel | 6x the exports an hour | nothing | +| Rig on the installer scripts | a prepare miss costs a restart per 40 min instead of per hour | a miss is 10% of an epoch | prepare-ahead as the only path (section 9) | +| Pool user | nothing | nothing | nothing | +| Every node's CPU (the VDF) | one core 25% busy | one core 100% busy | peers' proofs for a node with no core to spare (section 3.3) | +| The chain | difficulty settles 144 s per epoch: 6% | 24% | the +-15% settle measurement (section 9) | + +What the same step costs the adversary, from section 5.1's rule `max(0, 1 - (compile - 600) / epoch_len)`: + +| Compile time per program | Share of each epoch mined at 3,600 | at 2,400 (first step) | at 1,800 (second step) | at 600 (the floor) | +|---|---|---|---|---| +| 12 min (PRflow partitioned, a small kernel) | 97% | 95% | 93% | 0% | +| 42 min (PRflow monolithic) | 47% | 20% | 0% | 0% | +| 160 min (PRflow worst) | 0% | 0% | 0% | 0% | + +Reading: the first step cuts the 42-minute class from half of every epoch to a fifth and the second step removes it; the 12-minute class (a small kernel on a fast flow) survives every step but the floor, so an alert that persists through two steps is the case the floor was reserved for. A soft overlay compiles nothing and is untouched by every step: that lane is section 12. + +## 12. The FPGA lane and the ranking against layer 7 (Counter ASIC 3.0 item 5) + +6 October 2026, worker ca3-detector. No hardware was measured here; every FPGA figure is a product figure or a published measurement with its source and date read, and every derived number is labelled. The GPU side is the measured RTX 5090: 17.5 G dependent 4-byte reads per second at 415 ns (`docs/benchmarks/repro.md` section 2.2, 6 October 2026) at about 326 W (the bench log's 328.6 W peak, approximate). + +### 12.1 The parts + +| Part | HBM | Bandwidth | Stacks, pseudo-channels | Board power | Source (read 6 October 2026) | +|---|---|---|---|---|---| +| AMD Alveo U55C (XCU55, Virtex UltraScale+) | 16 GB HBM2 | 460 GB/s | 2 stacks; 32 pseudo-channels ("32 independent pseudo-channels") | 150 W maximum total, 115 W typical (TDP) | https://www.amd.com/en/products/accelerators/alveo/u55c/a-u55c-p00g-pq-g.html (through a search summary, the page itself timed out); the CoreEL data sheet https://www.c2s.gov.in/Technical_Data_sheet/new/Alveo_U55C_data_sheet.pdf ("HBM Memory 16GB", "HBM Bandwidth 460 GB/s", "Power (TDP) 115W", 1,304K LUTs, 9,024 DSP slices) | +| AMD Alveo U280 (XCU280) | 8 GB HBM2 | 460 GB/s | 2 stacks of 4 GB, each 8 channels of 2 pseudo-channels: 32 pseudo-channels, 32 AXI ports at up to 450 MHz | 225 W total electrical card load | DS963 v1.3 (May 2020) through https://www.digikey.com/en/htmldatasheets/production/3778633/0/0/1/a-u280-a32g-dev-g ; the channel layout from Shuhai (Wang, Huang, Alonso, FCCM 2020, https://arxiv.org/pdf/2005.04324, section II.A and III) | +| AMD Versal HBM (VH1582 class) | 32 GB HBM2e | 819 GB/s | 2 stacks (approximate: the series page gives capacity and bandwidth, not the stack count) | not a board; the VHK158 evaluation kit's power is not published as a product figure (approximate: 150 to 250 W for a card around it) | https://www.amd.com/en/products/adaptive-socs-and-fpgas/versal/hbm-series.html through a search summary ("819 GB/s of memory bandwidth and 32 GB of capacity") | +| Intel (Altera) Agilex 7 M-series | up to 32 GB HBM2e | 820 GB/s (410 per stack) | 2 stacks | not published as a board figure | the Agilex 7 M-series memory-bandwidth white paper, https://www.intel.com/content/dam/www/central-libraries/us/en/documents/2022-12/agilex-7-fpgas-m-series-memory-bandwidth-white-paper.pdf , through a search summary; the product page redirected to a 404 on 6 October 2026 | + +### 12.2 Random reads in flight per watt + +The metric of the plan: reads in flight = (random reads per second the HBM controller sustains) x (its random-access latency), divided by board watts; the reads per second per watt column is the same quantity without the latency factor and is the one that sets hash rate per watt. + +| Figure | Value | Source and label | +|---|---|---| +| HBM2 idle read latency on the U280, from the FPGA fabric | page hit 106.7 ns, page closed 122.2 ns, page miss 137.8 ns (48 / 55 / 62 cycles at 450 MHz) | Shuhai, Table IV (measured) | +| HBM2 random-access throughput as measured on the U280 at the default address mapping (one bank active per channel, RGBCG) | 2.4 GB/s per AXI channel at 32-byte bursts, 4 KB stride, 256 MB working set = 75 M random reads per second per pseudo-channel, 2.4 G per card over 32 | Shuhai, section IV.C and Figure 7 (measured; the paper's point is that this mapping is the wrong one for random access) | +| HBM2 random-access ceiling, bank-bound, with a bank-interleaved mapping | banks per pseudo-channel / row cycle: 8 to 16 banks / 45 ns = 178 to 356 M activates per second per pseudo-channel, 5.7 to 11.4 G per card over 32 | approximate: tRC 45 ns from JEDEC HBM2 as quoted at https://www.overclock.net/threads/the-hbm2-timings-thread.1743510/ and about 48 ns in the MEMSYS 2018 paper the history cites ([L1]); the banks per pseudo-channel from memory; the tFAW and tRRD command limits are not applied, which makes this a ceiling | +| Queue depth from the fabric | 32 AXI ports; outstanding reads per port in the AMD HBM IP (PG276) not read today | approximate: at 64 per port the fabric holds 2,048 reads in flight, which at 137.8 ns is 14.9 G per second, above the bank ceiling, so the banks bind, not the queues | +| HBM2e parts (Versal HBM, Agilex M) | bandwidth 1.8x the U55C's; the row cycle is the same DRAM | the history's [L1] reading: HBM raises bandwidth, not the row cycle; so the random-read ceiling per stack is the HBM2 ceiling within the clock ratio (approximate) | + +| Design | Random reads per second | Latency | Reads in flight | Board W | Reads in flight per W | Reads per second per W | Against the 5090 (per W) | +|---|---|---|---|---|---|---|---| +| RTX 5090, measured | 17.5 G | 415 ns | 7,260 | 326 (approximate) | 22.3 | 53.7 M | 1.0x | +| U55C overlay at Shuhai's measured random rate (default mapping) | 2.4 G | 137.8 ns | 330 | 115 to 150 | 2.2 to 2.9 | 16 to 21 M | 0.30x to 0.39x reads per second per W; 0.10x to 0.13x in flight per W | +| U55C overlay at the bank-bound ceiling (bank-interleaved mapping, approximate) | 5.7 to 11.4 G | 137.8 ns | 790 to 1,570 | 115 to 150 | 5.2 to 13.7 | 38 to 99 M | 0.71x to 1.85x reads per second per W; 0.23x to 0.61x in flight per W | +| U280 overlay, the same ceilings | as the U55C (the same HBM subsystem) | 137.8 ns | the same | 225 | 1.5 to 7.0 | 11 to 51 M | 0.20x to 0.94x reads per second per W | +| Versal HBM or Agilex M overlay (HBM2e, 2 stacks) | the HBM2 ceilings times at most the clock ratio 1.8 (approximate) | about the same | | approximate 150 to 250 | | | under 2x at the ceiling, under 1x at Shuhai's measured rate | + +Reading, with every caveat in the table: at the only measured FPGA random-read rate in the literature found today (Shuhai's 2.4 G per second for a two-stack HBM2 card) a soft-overlay FPGA sits at a third of the 5090 per watt, in the RX 9070 XT's class (2.4 to 2.5 G per second, `docs/bench-log.md`); at the bank-bound ceiling that no published design reaches it could sit between 0.7x and 1.9x per watt, which is why the ceiling row is kept and labelled. The overlay's ALU side is not the bound: a hash is 64 instructions x 8 iterations per lane, 70 G instructions per second at the 5090's 137 MH/s, about 175 soft 32-bit ALUs at 400 MHz on a 1.3 M-LUT part (approximate), and the mixer is not on the hash path. What is owed to turn the ceiling row into a number: one HBM FPGA under a bank-interleaved random-read kernel (a rented U55C or U280 hour, the chase kernel of `docs/benchmarks/repro.md` ported to an AXI master), measured in reads per second and watts; until then the public claim carries the measured row (0.3x to 0.4x per watt) and names the ceiling. + +### 12.3 Compile-ahead: the two FPGA lanes in one table + +| Lane | Compile per program | Share of a 600-s epoch it can mine (section 5.1) | Share of a 3,600-s epoch | Rate against the 5090 per watt | What answers it | +|---|---|---|---|---|---| +| Hard datapath (a bitstream per program) | 42 to 160 min monolithic, 12 to 18 min partitioned (PRflow, FPT 2019, section 5.1) | 0% | 47% at 42 min, 97% at 12 min | a hard datapath has no overlay overhead, but it still reads the same HBM: bounded by the same 0.3x to 1.9x per watt, before its 0% to 97% duty | layer 9: the epoch length (section 11) | +| Soft overlay (a processor on the fabric; the program is data) | none: 0% of every epoch | 100% | 100% | 0.3x to 0.4x per watt measured-basis, up to 1.9x at the unproven bank ceiling | the latency bound itself (the reads per watt), the dataset and cache sizes, the mixer; not the epoch length | + +### 12.4 Layer 9 (epoch length) against layer 7 (the mm8 reserve family): the ranking + +| Adversary class | What layer 9 costs it | What layer 7 costs it | +|---|---|---| +| Hard-datapath FPGA | its whole duty at the floor (0% of every epoch; 20% at the first step for the 42-min class, section 11) | nothing: it already compiles a datapath per program, and a new family is one more block in that datapath | +| Soft-overlay FPGA with HBM | nothing (no compile) | nearly nothing: FPGA DSP blocks do int8 dot products natively (the DSP58 of Versal and the DSP48E2 of UltraScale+ carry int8 multiply-accumulate; approximate, from memory), so an mm8 unit is cheap on the fabric | +| On-die 256 MiB recompute chip (`chip-model-v3.md`, 0.92x with the factor) | nothing (it executes the program) | little: int8 matrix blocks are licensable IP at every node (history section 4.3, addition 6) | +| Partial-store chip with a custom memory system (item 1) | nothing | little, as above | +| One-vendor GPU fleet (the 7.5x AMD gap, item 7) | nothing | it moves the gap: dp4a is 1.17x a step on the 5090, 1.06x on the 9070 XT, 1.6x emulated on Apple (`docs/bench-log.md`, layer 7 row) | + +| User tier | What layer 9 costs it (sections 7 and 11) | What layer 7 costs it (the bench log's layer 7 row, `int8-matrix-family.md`) | +|---|---|---| +| Home miner, one NVIDIA card (8 to 32 GB), Windows or Linux | 0.04% at 2,400, 0.2% at 600 | nothing relative: dp4a is native, 1.17x a step | +| Home miner, one AMD RDNA 4 card | compile owed; 0.31 s plus it | dp4a native at 1.06x a step: falls 10% behind NVIDIA per step, relatively | +| Home miner, one Apple card, macOS | 0.02% at 2,400 with the race off; 1.6% with it on | the step is emulated at 1.6x the cost: the Apple tier loses share to every native card for as long as the family is live | +| Integrated GPU (AMD, Intel) | 0.3 to 0.5% at 2,400; 5.2% under load until per-day dataset reuse | as the vendor's discrete part (native on AMD, emulated on Intel: not measured) | +| Rig | exports per epoch, parallel prepares | as its cards | +| Pool user | nothing | nothing directly; the pool's card mix shifts | +| Every node (the VDF) | one core 25% busy at 2,400, 100% at 600 | nothing | +| The chain | 6% of each epoch in difficulty settle at 2,400, 24% at 600 | nothing | +| Reversibility | by the same 90% signal, up or down, in 9 days | a family unlocked by height stays; by signal it can be voted off, but the vendors' relative rates are what they are while it is on | + +Ranking: layer 9 sits above layer 7 in the reserve. Layer 9 removes one adversary class outright (the per-program bitstream, the first adversary of Lyra2REv2 and X16R in the history), costs every GPU tier under 2% at the first step and under 1% with the race off, and is reversible by the same signal that set it. Layer 7 removes no adversary class (every chip and every FPGA has an int8 dot product; the history's "watch ML hardware"), and its cost lands on one honest tier, Apple at 1.6x per emulated op, with a 10% relative shift against AMD. The chip model does not move for either (the on-die recompute chip's cost is the item derivation): layer 9's value is response time against the FPGA lane, layer 7's is a family a 12-op chip lacks, which the overlay and the chip both acquire cheaply. Item 6 (order the reserve by chip-unfriendliness, mm8 last) follows from the same two tables. Decision for Josh: reserve order layer 9 (epoch length, already live at the base) first, the 32-bit datapath families next, mm8 last; and fund one HBM FPGA hour to replace the ceiling row of 12.2 with a measurement before the public testnet's benchmark page claims a number for this lane. diff --git a/docs/plans/funding.md b/docs/plans/funding.md index ad1f0d4bf..793dccb4c 100644 --- a/docs/plans/funding.md +++ b/docs/plans/funding.md @@ -62,3 +62,12 @@ The fee is 1% of rewards on the official client. Rewards in year one are 963 mil 2. A review or audit that is paid for is published whole, pass or fail, and linked from `docs/evidence.md`. 3. A bounty is announced only when it is escrowed. 4. This plan is revised when a number changes; the git history of this file is the record. +5. The chip bounty's trigger is daily issuance in dollars, not a date (Counter ASIC 3.0 item 4b, 6 October 2026): **the bounty is escrowed and the benchmark page is live before daily issuance crosses USD 20,000 a day.** Daily issuance is blocks per day times the subsidy (spec 2.5, `site/lib/emission.mjs`: 3,168,808,781 sompi per DAA second in period 0, so 2,737,851 IGN a day after the 30-day ramp, 273,785 on day 0; 1,368,925 a day in period 1, years 3 and 4), times the price. The history (`docs/analysis/asic-resistance-history.md` section 2.5) puts the first public chip on compute-bound hashes at USD 21,000 to 31,000 of daily issuance (Kadena, Radiant, Handshake) and Vorick's 2018 rule at about USD 55,000 a day; USD 20,000 sits under the lowest observed arrival, so the escrow lands before any chain in that table got its chip. The operating entity watches the number (owed: an "issuance per day in dollars against the USD 20,000 line" row in the 08:00 daily report) and the detector (`tools/observer/detector.mjs`) runs from the public testnet, where issuance in dollars is zero and the clock has not started. The prices at which the line is crossed, so the number is concrete: + +| Daily issuance line | Period 0 (year 1 to 2, after the ramp): price per IGN | Period 1 (years 3 to 4) | Period 2 (years 5 to 6) | +|---|---|---|---| +| USD 20,000 (the rule) | USD 0.0073 | USD 0.0146 | USD 0.0292 | +| USD 30,000 (the top of the compute-bound arrivals) | USD 0.0110 | USD 0.0219 | USD 0.0438 | +| USD 55,000 (Vorick) | USD 0.0201 | USD 0.0402 | USD 0.0804 | + + What it means per tier: nothing changes in the protocol at the line; a home miner on any card can read the live benchmark page and the bounty terms, so a chip's existence becomes something its designer is paid to disclose rather than to hide; a pool user sees the same page. If the entity cannot fund the escrow when the line approaches, rule 3 holds (nothing is announced) and the detector plus the epoch-length signal (`docs/plans/epoch-length.md` section 11) are the response that costs no money.