From a59e782159d4fe1a5d8dd52fa3cc94d242d0f55d Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Mon, 5 Oct 2026 22:20:05 +0000 Subject: [PATCH] Bench log: Counter ASIC 2.0, the numbers (the level 3 page section); litepaper anchor --- docs/bench-log.md | 32 ++++++++++++++++++++++++++++++++ site/litepaper.html | 2 +- 2 files changed, 33 insertions(+), 1 deletion(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index 5b674a242..bc0cfe101 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1586,3 +1586,35 @@ Machine: PC 1 (machine ae432dc7), Windows 11, WSL2 Ubuntu 24.04 as root, 16 core For comparison (this log): the Apple M5 Max CPU on 4 October, loaded, block-56 shard 0: core 83.1 s, compressed 272.3 s; on 3 October the v0 guest on block-78: core 22.0 s, compressed 55.7 s. The RTX 5090: block-78 core 1.4 s, compressed 2.7 s (4 October, mining paused); a full shard at `S_p` compressed 10.9 s alone and 33.0 s beside the miner; an empty live shard 7.0 to 7.7 s beside the miner (5 October). No fresh Mac run tonight: the measure lock was held from 20:31Z (a read-width `packbench`, three builds, a 1,500-s proving-v1 network under `run`) and did not free inside the 10-minute window set for it. Reading, and the consequences (CLAUDE.md, every number). Doubling the cycles added 4.5 s to the core proof and 3.1 s to the compressed proof: about 280 s of a CPU proof is fixed cost in the compressed-proof recursion, so no shard size brings a CPU proof under the launch deadline (20 to 60 s behind the tip) or near the 10-s assignment window; it fits only the v1 unproven deadline (600 s), which pays a CPU prover only when no card has proven the shard in 10 minutes. The 29.5 to 30.5 GB peak RSS means the CPU prover needs 32 GB free: a 64 GB Windows PC (WSL2 takes half the host's RAM by default), a 32 GB Linux machine, a 64 GB Mac; a 16 GB machine cannot run it at all. Per tier: an AMD-only home miner (8, 12 or 16 GB, Windows or Linux) mines and does not prove, and loses the 20% proving-pool share; Apple silicon the same (the M5 Max mines at 26.7 MH/s, this log, 4 October); a mixed rig proves on its NVIDIA cards and the rig installer's `prover_decision` already skips every non-NVIDIA card (`packaging/linux/bin/igneum-rig-lib.sh`, branch `rig-install`), now a stated requirement; the app's `provedefault.rs` already keeps proving off on Apple silicon and off without an NVIDIA card. Decision asked of nobody: no CPU tier (the analysis, section 4a); the public line for the site, litepaper and Proving tile is in section 4c ("Proving needs an NVIDIA card with 16 GB or more today ... AMD and Apple cards mine. A prover for them lands when a zkVM ships one"). The first job proved nothing because an apostrophe inside a single-quoted awk program ended the quote and bash refused the loop while the job reported exit 0; the class fix is `tools/amd-prove/check-job-bash.sh` (`bash -n` on the embedded bash body before publishing) and the same `bash -n` inside the job before the run, both shown to refuse the bad body and pass the fixed one. + +## Counter ASIC 2.0, the numbers + +5 October 2026 (night). The chip-resistance layers measured on the three cards we own (Apple M5 Max, RTX 5090 on PC 1 and PC 2, RX 9070 XT on PC 1's eGPU), the decisions taken under the project lead's delegated rules for the devnet, and the chip model before and after. Every number is from an entry above or from the plan documents named; approximate is marked. Levels: `docs/plans/counter-asic-2-public.md`. + +**Program class v3 (the devnet, activation by height switch `program_class_v3_activation_daa`)** = class v2's 128 x 4-byte loads, the era draw of the table layout and the working-set windows (layers 4 and 8), the cache growth rule (layer 6, option C: the cache doubles when the dataset doubles), the mixer at x8 (M16's multiplier), reserve family R1 (integer matrix, switched off) and the epoch length as a signalled reserve parameter (layer 9, 3,600 DAA s until a 90% signal). Not adopted on the measurements: wider reads (layer 1), the per-load width mix (layer 2), the per-warp write scratch (layer 3), the hot table (layer 5). + +| Card | v2 MH/s | v3 MH/s, six eras (spread) | Bytes per hash | Latency-bound share | Daily 1 GiB build, v2 / v3 | +|---|---|---|---|---|---| +| Apple M5 Max, Metal | 27.68 | 28.35 to 28.58 (0.8%) | 512 | 1.06 | 21 / 21 ms | +| RTX 5090, CUDA | 137.2 | 136.18 to 138.01 (1.3%) | 512 | 1.01 | 25 / 23 ms | +| RX 9070 XT, OpenCL | 18.09 | 18.61 to 19.21 (3.2%) | 512 | 0.95 | 74 / 75 ms | + +CPU verifier, one M5 Max core (loaded box; ratios are the measurement): v2 0.61 ms per warp, v3 (x8) 2.08 ms, worst cold 2.15; the 10 ms gate holds 4.8x. Bit-exact: every v3 pack's fingerprint equal on Metal, Apple OpenCL, CUDA and AMD OpenCL. + +| Layer | Measured | Decision | The number | +|---|---|---|---| +| 1 wider reads | w16 139.8 / 17.90 / 28.26 MH/s (5090 / 9070 XT / M5 Max) against v2 136.1 / 18.15 / 27.74; w64 71.9 on the 5090 (share 0.58, 37% of its stream) | out: keep 4 B | the 9070 XT does 2.4 G dependent reads/s at every width; wider reads make the 5090 bandwidth-bound | +| 2 width mix per load | spread over six programs 18.8 / 7.4 / 11.3% and 22.3 / 5.5 / 8.1% | out | the 5% rule | +| 3 write scratch | GPU cost 12 to 48% at 32 and 128 KB per warp; the on-die-cache chip 2.4x at every share | out (the construct is sound; its tests stay) | the verifier resets the scratch per unit, so a chip keeps it in 80 to 320 B per lane | +| 4 and 8 era layout and windows | six-era spread 1.3 / 3.2 / 0.8% | in | under the 5% rule; the SRAM mirror a chip needs is the whole dataset every hour | +| 5 hot table | added form g 0.87 / 0.85 / 0.84 (5090), 0.84 / 0.81 / 0.80 (9070 XT) at 32 / 64 / 96 MiB | out (a 3.0 option) | no card keeps 32 MiB resident while the dataset streams; the replaced form helps the chip | +| 6 cache schedule | the 256 MiB mirror is 128 mm^2 and $46 at N5 by shipped cache-die density, approximate | option C, in | the cache's job is to stay above GPU L2 (96 MB on the 5090, 128 MB on GB202) | +| 7 integer matrix | dp4a 1.17x a step on the 5090, 1.06x on the 9070 XT, 1.6x emulated on Apple; mm8 native on all three as a tile | reserved R1, off | unlock at era 4 or 90% signal | +| M16 mixer | x4: verifier 1.24 ms, chip 1.84x with the allowance; x8: 2.08 ms, 0.92x; the daily build unmoved on every card | x8 in | the only lever that moves the named chip | +| 9 epoch length | compile-ahead 0.5 s (M5 Max, race off), 1.0 s (5090), 38 s with the race; FPGA compiles 42 to 160 min (PRflow, FPT 2019) | reserved, 600 s to 2 h by signal | at 600 s a per-program bitstream mines 0% of each epoch | + +**The chip model, before and after** (`docs/analysis/chip-model-v3.md`, `docs/analysis/sram-mirror.md`): the strongest chip we can name holds the whole 256 MiB cache on-die (about 128 mm^2 and $46 of silicon at N5, approximate) and computes dataset items on the fly at 50 T integer op/s. Against the RTX 5090's measured 136.1 MH/s: class v2 333 MH/s, 2.4x; class v3 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim is "under 2x"; the margin is thin on the allowance (3.3x reads 1.0x) and 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset. + +**The user tiers.** AMD RDNA 4 sits at about a seventh of a 5090 on this hash (its dependent-read rate: 2.4 G against 17.5 G per second), 2.2x worse per pound at list prices and 4.9x worse per watt (approximate); the card's memory system, not a tuning gap. The integrated tier on the CUDA and OpenCL one-click workers mines v3 with a restart per epoch until per-day dataset reuse lands (0.3.12). Card lifetime under the step schedule: a 4 GB card to year 4, 8 GB to year 12, 12 GB to year 28 with the cache freed after the daily build. + +**The bounty.** A standing bounty for any chip design beating a GPU by more than 2x on the published model, with a leaderboard by card model, January 2027 (spec O-1.17). diff --git a/site/litepaper.html b/site/litepaper.html index bc7e21beb..73aaa2049 100644 --- a/site/litepaper.html +++ b/site/litepaper.html @@ -421,7 +421,7 @@ body.all .pager{display:none}

Three ideas carry the chip resistance. The hash rewrites itself. A new program every hour, drawn from the chain. Its memory pattern changes with it. The rules change on a schedule fixed at launch. No release, no vote. It waits on memory, not maths. Every hash is a chain of random reads into a table too big for a chip to carry. The wait is the same physics for everyone. Miners hold the switch. Spare defences are written into the rules, switched off. A 90% miner signal turns one on. No fork.

-

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model and the bounty are public: the numbers. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

+

No hash has stayed free of chips forever. Igneum does not claim to. It claims the gain is small, the response takes a week, and both are measured. The model and the bounty are public: the numbers. Monero has run on RandomX since 2019 with no chip publicly shipped, approximate; that is precedent, not proof.

One thing takes a person, here and on every chain that exists: writing new code. A chain cannot safely write its own generator, and it cannot safely tell a chip from a wave of honest new cards by hashrate alone. If the design above ever failed, anyone could publish a new generator and miners would switch it on by signalling, as Monero's community can fork. Igneum is built to make that day unlikely, and does not depend on avoiding it.