Site and litepaper: level 1 on the hero and abstract, the chip bullet on the measured row; the level 3 table filled; status 22:18

This commit is contained in:
igneum-labs 2026-10-05 22:18:39 +00:00
parent 5959651452
commit 9488399fcc
4 changed files with 19 additions and 9 deletions

View file

@ -28,13 +28,19 @@ Headline of the chip model (5 October 2026, night): the strongest chip holds the
Per card, the bench table: the v2 class and the v3 class, hash rate, bytes per hash, the latency-bound share (rate over the card's random-read ceiling per load), the CPU verifier per warp, with machine, date and command. The chip model before and after Counter ASIC 2.0 (the m16 model's gain arithmetic at the v2 class and at the v3 class, with the SRAM a mirror needs, cited or approximate as the analysis says). The bounty terms (spec O-1.17: the leaderboard by card model, the standing bounty for any chip design beating a GPU by more than 2x, January 2027). Here the layers are named next to their numbers: read width, per-program mix, scratch, era layout, working set, hot table, cache schedule, the reserved integer-matrix family.
| Card | v2 MH/s | v3 MH/s | v3 bytes per hash | Latency-bound share v3 | Verifier ms per warp v3 |
|---|---|---|---|---|---|
| Apple M5 Max (Metal) | 27.74 | [owed: the v3 pack] | 512 + hot | [owed] | [owed: ca2-mixer] |
| RTX 5090 (CUDA) | 136.1 | [owed] | 512 + hot | 0.96 at v2 | [owed] |
| RX 9070 XT (OpenCL, eGPU) | 18.15 | [owed] | 512 + hot | 0.87 at v2 | [owed] |
| Card | v2 MH/s | v3 MH/s (era packs, six eras) | Bytes per hash | Latency-bound share | Verifier ms per warp (v2 / v3, one loaded M5 Max core) | Daily 1 GiB build (v2 / v3) |
|---|---|---|---|---|---|---|
| Apple M5 Max (Metal) | 27.68 | 28.35 to 28.58 (spread 0.8%) | 512 | 1.06 | 0.61 / 2.08 | 21 / 21 ms |
| RTX 5090 (CUDA) | 137.2 | 136.18 to 138.01 (spread 1.3%) | 512 | 1.01 | the same verifier | 25 / 23 ms |
| RX 9070 XT (OpenCL) | 18.09 | 18.61 to 19.21 (spread 3.2%) | 512 | 0.95 | the same verifier | 74 / 75 ms |
The integrated tier (Radeon iGPU, Intel UHD) on the CUDA and OpenCL one-click workers mines v3 with a restart per hourly epoch, because those workers rebuild the dataset on every prepare (7 to 12 s at x1, about 4x at the x4 mixer; the Metal worker keeps its day's dataset); per-day dataset reuse in those two workers is the first item after the publish (0.3.12). AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap.
Every number measured 5 October 2026 (`docs/plans/era-layout.md`, `docs/plans/mixer-x4.md`, `docs/bench-log.md`); the v3 verifier figure is the x8 mixer on a loaded core (about 1.3 ms quiet, approximate). Bit-exact: every v3 pack's fingerprint equal on the three vendors.
Chip model, before and after (`docs/analysis/chip-model-v3.md`): the on-die-cache recompute chip (the whole 256 MiB cache in SRAM, about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density, approximate) against the RTX 5090's measured 136.1 MH/s at 50 T integer op/s: class v2 333 MH/s, 2.4x; class v3 (mixer x8) 41.7 MH/s, 0.31x bare, 0.92x with a 3x fixed-function allowance (approximate), 0.76x at equal silicon. The claim "under 2x" holds with the margin stated: 8% on the allowance (a 3.3x allowance reads 1.0x), 9% on the budget. Next levers, named: the mixer at x16 (the verifier at about 4 ms per warp, inside the 10 ms gate; a 2019-class core unmeasured), a hot table small enough to stay resident beside the streaming dataset (measured and not adopted tonight: 32 to 96 MiB tables cost the GPU 7 to 20% and help the chip).
Levers measured and not adopted (5 October 2026): wider reads (16 and 64 B: no card gains, the 5090 goes bandwidth-bound at 64 B), the per-load width mix (5.5 to 22.3% spread), a per-warp write scratch (the chip keeps it implicitly: 2.4x at every share), the hot table (above). Reserved, switched off: the integer matrix family R1 (mm8, native on all three vendors as a tile; dp4a 1.17x a step on the 5090, 1.06x on the 9070 XT, emulation 1.6x on Apple) and the epoch length (600 s to 2 hours by 90% signal; a per-program FPGA bitstream mines 0% of a 600-s epoch).
AMD RDNA 4 sits at about a seventh of a 5090 on this hash by its dependent-read rate (2.4 G against 17.5 G reads per second), 2.2x worse per pound at list prices and 4.9x worse per watt (read-width.md section 4.1, approximate); the card's memory system, not a tuning gap.
Chip model, before and after: [owed: from docs/analysis/sram-mirror.md after the shipped-density correction (256 MiB on-die at about 130 to 165 mm^2 by AMD 3D V-Cache and TSMC N5 macro density, 54 to 83 mm^2 bit-cell-only lower bound), the hot-table and scratch analyses; the on-die-cache recompute chip is a named row per variant].

View file

@ -517,3 +517,7 @@ The card draws 302 to 316 W under this program whatever the cap, so a cap above
## 22:16 gate G6 GREEN on the final tree
build-20261005-221237 (main d233fa1, fork 89dfcb95, 185 s): every stage ok; kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test, with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8; 0 failed; with job 2's kaspa-consensus 97 alone, G6 is green. PC 2 goes to the prover-floor agent's 90-minute build (its go at 22:17), then the aggregation-cost re-run, then the prover-floor sweep. Gates: G3 green (the Mac suites on the final class: 53 + 4 + 19 + 7 crate tests, the Metal fuzz, edge, stats and determinism runs, the scratch tests), G4 green (runs 1 and 2; run 3 on the final x8 + era class pending), G6 green; G1 and G2 pending the era agent's PC 1 job (running from 22:16); G5 (the Windows and Mac workers from the same commit) is the ship's build step on the merged tree.
## 22:18 the spec and the public copy carry the final class
docs/spec/01-lottery-hash.md on ca2-coord: 1.8.5 (the mixer x8 form with the measured costs), 1.13.1 (the era draw: stride, interleave, windows, the devnet stand-in, the measured spread), 1.17 (the class v3 vectors), 1.5 (the cache note); earlier tonight 1.12 and 1.13.1 (epoch_len), 1.13.2 (R1 and the emulation rule), 1.13.3 (option C and the step mapping), 1.4.5 and 1.4.6 (generator 3, the class and era in the pack), and 4.3. Public copy: level 1 on the hero and the abstract ("Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public."), the limits bullet rewritten on the chip row (0.92x with the allowance, approximate; the margin on the numbers page; the next lever named), the level 3 table filled from the final numbers in counter-asic-2-public.md (the bench page section is written from it at the ship). Waiting: the era agent's PC 1 job (G1 on the final class and G2), the node agent's gate run 3; then the integration merge and the ship.

View file

@ -340,7 +340,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var(
<div class="hero-copy">
<div class="eyebrow ember">GPUs are back · for good</div>
<h1>Mined by GPUs.<br>Proven by fire.</h1>
<p class="lead">A chain built so a chip gains too little to take your place. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.</p>
<p class="lead">Built for graphics cards. A custom chip gains under 2x, and the model and the bounty are public. A mining program that rewrites itself every hour, so a chip built for one hour is useless the next. An NVIDIA card with 24 GB or more proves every block and gets paid for it; AMD and Apple cards mine. No premine, no stake, no foundation, no merge to proof of stake.</p>
<div class="cta">
<a href="#mine" class="btn primary">See the miner</a>
<a href="/litepaper" class="btn">Read the litepaper</a>

View file

@ -298,7 +298,7 @@ body.all .pager{display:none}
<article>
<section id="abstract">
<h2>Abstract</h2>
<p class="lead">Igneum is a proof-of-work blockchain mined on graphics cards, where the same cards prove every block with zero-knowledge proofs and sell proving to other chains.</p>
<p class="lead">Igneum is a proof-of-work blockchain built for graphics cards, where NVIDIA cards also prove every block with zero-knowledge proofs and sell proving to other chains. A custom chip gains under 2x, and the model and the bounty are public.</p>
<p>It runs the Ethereum virtual machine, so anything built for Ethereum runs on Igneum unchanged. Transactions are included in about one second, proven within about a minute at launch, and locked by miners within about two. There is no premine, no pre-sale, no treasury taken from emission, no stake anywhere in consensus, and no dependence on any other chain. Mining stays open to anyone with a GPU because the mining program changes every hour, so a chip built for one program is useless for the next, and a chip for the whole program space is a GPU without the graphics parts. No scheduled human release is needed to keep it that way. Writing new code, including an emergency fix to the proof system, is the one thing that takes a person, and it activates only on miner signalling.</p>
<div class="stats">
<div class="stat"><div class="v">1 / s</div><div class="k">blocks, rising to 10</div></div>
@ -730,7 +730,7 @@ body.all .pager{display:none}
<p>Here are the limits, stated before anyone else states them.</p>
<ul>
<li><strong>A proof in seconds.</strong> Not at launch. Proving a full block today needs a cluster of 100 to 200 consumer GPUs, approximate, so Igneum launches with proofs within about a minute and tightens as hardware improves. Users still see their transaction land in one second.</li>
<li><strong>A chip is impossible.</strong> No. A chip is a bad bet, because the target moves before it ships. The efficiency ceiling for a fixed chip on a memory-bound program is a target of under 2x, not a measurement, and Igneum's generator changes under it every hour. Monero's seven years without a public chip are precedent, not proof.</li>
<li><strong>A chip is impossible.</strong> No. A chip is a bad bet, because the target moves before it ships. The published model (5 October 2026) prices the strongest chip we can name, one with the whole cache on-die computing dataset items on the fly, at 0.92x the hash rate of an RTX 5090 per unit of silicon with a 3x fixed-function allowance, approximate; the claim is under 2x, the margin is stated on the numbers page, and the next lever is named there. No hash has stayed free of chips forever; Igneum does not claim to. Monero's seven years without a public chip are precedent, not proof.</li>
<li><strong>A guaranteed income floor.</strong> No. External proving is a small market today. Igneum's miners' marginal cost in it is close to power, which is an edge and nothing more.</li>
<li><strong>A memory-hard prototype on every vendor.</strong> Not yet. The 256 MB cache closed the shortcut on Apple silicon (computing items runs 4.8x slower than loading them, measured 3 October 2026). The same ratio on NVIDIA and on a discrete AMD card is Open.</li>
<li><strong>Finality in the first month.</strong> No. No checkpoint locks until the 30-day window has 30 days of history. The first month of mainnet is proof of work with a 12-hour depth, and the text above says so wherever a day count appears.</li>