Merge counter-asic-4 91929389 into master (gate: green on 91929389, recorded by tools/ci/pre-push.sh; landed on the build mirror)

This commit is contained in:
igneum-labs 2026-10-08 16:03:13 +00:00
commit ba6bf0ea16
4 changed files with 68 additions and 6 deletions

View file

@ -89,7 +89,7 @@ Reading: layer 2 does not move the chip anyone builds, because a chip buys DRAM
### 3.3 Per tier
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). **The DRAM-read cost per hash on the NVIDIA cards is NOT size-independent at the knee, MEASURED (the hash lane's kit b, PC 1, 12:49 to 12:58 UK, the pinned class v3 program 73bcbfe8 at 2, 4 and 8 GiB against the 1 GiB control 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at the 1,300 MHz lock; 250 batches per row, fingerprints PASS): unlocked 133.86 at 315.7 W (-2.8 percent), 132.42 at 317.2 (-3.8), 131.75 at 319.2 (-4.3); at the lock 121.14 at 210.3 W (-4.9 percent), 113.06 at 204.5 (-11.2), 109.39 at 201.6 (-14.1); MH/W at the lock 0.599, 0.576, 0.553, 0.543, which is 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB.** The lane's earlier reading (2 MiB pages keep the TLB's reach past 8 GiB) holds unlocked, where the card hides most of the page-walk term in its slack; the latency-bound regime at the lock exposes it. **The genesis floor itself, 5.5 GiB, MEASURED on a rented 5090 at stock (the fleet hand, RunPod secure, driver 570.195, 15:19 to 15:23 UK, the ca3-ds55 kit's own worker, program 73bcbfe8 in both packs, 250 and 500 x 2^24, every row PASS with 96 of 96 vector lanes): the 1 GiB control 141.48 MH/s at 325.6 W (2.30 microjoules; memory.used peak 1,914 MiB) against ds55 136.56 at 305.3 W busy, 327.6 steady (2.24 busy, 2.40 on the steady watts; peak 6,522 MiB: the 5.9 GB dataset plus the 268 MB cache plus the context), fingerprints stable across both passes (ds55 23ced07a4d28b465 becomes the pin). So the 5.5 GiB floor costs 3.5 percent of the rate at the same watts, about 4 percent more energy per hash, and the non-power-of-two mapping (1,476,395,008 words, 92,274,688 items, loads as (src x words) >> 32) is not a cliff on sm_120. Caveat: this host's 5090 plateaued at 328 W on both packs (a host power cap; another host's 5090 pulled 443 W on class v5 genesis this afternoon), so the microjoules are capped-card numbers and the rate and fingerprints are the row.** Lane 3's reading of it for the schedule (2a61cb46, 15:25 UK): the step lands between the 2 and 4 GiB stock rows (2.8 and 3.8 percent), so the non-power-of-two floors of 5.5 / 8.5 / 11.5 cost nothing beyond the size and the multiply-shift mapping is safe to adopt at the v6 epoch; per tier the 5.5 GiB step costs a 5090 4 percent per hash at stock (measured) and about 9 at its knee (interpolated from the 2, 4, 8 GiB knee rows); the 8 GB tier's fate at that step is the RX 7600 reading from PC 1 (about 16:00 to 16:30 UK, an amendment; its 1.8 GB of headroom is the question), the 5.5 GiB knee row with it; the SRAM store at that step is 3 reticles, USD 1,500, its joules unmoved. So each step of the schedule costs a tuned 5090 about 4 to 5 percent per hash while the chip's joules do not move (floor lane 3: its ticket goes USD 1,500 / 2,500 / 3,000 at 5.5 / 8.5 / 11.5 GiB), and every chip edge against a card at its knee rises by 4 to 11 percent across the schedule; the honest sentence for the schedule decision is USD 1,000 of chip ticket per step for about 1 to 5 percent of the tuned 5090's energy and about a quarter of today's measured cards by count. **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). **The DRAM-read cost per hash on the NVIDIA cards is NOT size-independent at the knee, MEASURED (the hash lane's kit b, PC 1, 12:49 to 12:58 UK, the pinned class v3 program 73bcbfe8 at 2, 4 and 8 GiB against the 1 GiB control 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at the 1,300 MHz lock; 250 batches per row, fingerprints PASS): unlocked 133.86 at 315.7 W (-2.8 percent), 132.42 at 317.2 (-3.8), 131.75 at 319.2 (-4.3); at the lock 121.14 at 210.3 W (-4.9 percent), 113.06 at 204.5 (-11.2), 109.39 at 201.6 (-14.1); MH/W at the lock 0.599, 0.576, 0.553, 0.543, which is 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB.** The lane's earlier reading (2 MiB pages keep the TLB's reach past 8 GiB) holds unlocked, where the card hides most of the page-walk term in its slack; the latency-bound regime at the lock exposes it. **The genesis floor itself, 5.5 GiB, MEASURED on a rented 5090 at stock (the fleet hand, RunPod secure, driver 570.195, 15:19 to 15:23 UK, the ca3-ds55 kit's own worker, program 73bcbfe8 in both packs, 250 and 500 x 2^24, every row PASS with 96 of 96 vector lanes): the 1 GiB control 141.48 MH/s at 325.6 W (2.30 microjoules; memory.used peak 1,914 MiB) against ds55 136.56 at 305.3 W busy, 327.6 steady (2.24 busy, 2.40 on the steady watts; peak 6,522 MiB: the 5.9 GB dataset plus the 268 MB cache plus the context), fingerprints stable across both passes (ds55 23ced07a4d28b465 becomes the pin). So the 5.5 GiB floor costs 3.5 percent of the rate at the same watts, about 4 percent more energy per hash, and the non-power-of-two mapping (1,476,395,008 words, 92,274,688 items, loads as (src x words) >> 32) is not a cliff on sm_120. Caveat: this host's 5090 plateaued at 328 W on both packs (a host power cap; another host's 5090 pulled 443 W on class v5 genesis this afternoon), so the microjoules are capped-card numbers and the rate and fingerprints are the row.** Lane 3's reading of it for the schedule (2a61cb46, 15:25 UK): the step lands between the 2 and 4 GiB stock rows (2.8 and 3.8 percent), so the non-power-of-two floors of 5.5 / 8.5 / 11.5 cost nothing beyond the size and the multiply-shift mapping is safe to adopt at the v6 epoch; per tier the 5.5 GiB step costs a 5090 4 percent per hash at stock (measured) and about 9 at its knee (interpolated from the 2, 4, 8 GiB knee rows); the 8 GB tier holds at that step on an exact NVIDIA 8 GB card (the RTX 4060, 6,116 MiB resident, measured 16:38 UK, 10.0p) and on a 12 GB card (the RTX 3060, 6,129 MiB), the RX 7600 row for the AMD 8 GB case still owed, the 5.5 GiB knee row with it; the SRAM store at that step is 3 reticles, USD 1,500, its joules unmoved. So each step of the schedule costs a tuned 5090 about 4 to 5 percent per hash while the chip's joules do not move (floor lane 3: its ticket goes USD 1,500 / 2,500 / 3,000 at 5.5 / 8.5 / 11.5 GiB), and every chip edge against a card at its knee rises by 4 to 11 percent across the schedule; the honest sentence for the schedule decision is USD 1,000 of chip ticket per step for about 1 to 5 percent of the tuned 5090's energy and about a quarter of today's measured cards by count. **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
| Tier | At the floor (today to year 4) | At a 4 GiB state-driven step | At 16 GiB | Label |
|---|---|---|---|---|
@ -443,7 +443,7 @@ The conditions, read off the surface, which are the economic-resistance statemen
**The sentence, as the external review words it (10.0f item 5), served verbatim, with one word made honest (the site audit lane's read, 16:5x UK: class v4 and v5 have eight registers per lane and the window is class v6's new core shape, so "retains" is read as "retains across every rotation"):** "Class v6 adopts the 64-register window and retains it across every rotation. Current modelling estimates a 2.2x to 2.4x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.0x on the GPU's own node). The long-program and select-tree proposals were rejected. Economic resistance depends on development cost, deployment economics and productive hardware lifetime; family transitions receive an obsolescence benefit only where a loss of competitiveness is demonstrated; programmable multi-epoch designs are included in the assessment."
**Where the figures come from (the coordinator, 16:1x UK): the served sentence's numbers are read from the k lane's PLACED GATED rows (17:30 UK; 10.0i), not from 725d2945's close and not from the 16:0x synthesis; until they land the figures sit in 10.0i's bracket, and the serve slips to 18:30 if the rows are late.** The labels on its numbers as first written: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound.
**Where the figures come from (the coordinator, 16:1x UK): the served sentence's numbers are read from the k lane's PLACED GATED rows (10.0i), not from 725d2945's close and not from the 16:0x synthesis; the placement slipped to about 18:30 UK, so the sentence served at 18:30 carries 10.0i's tightened bracket, "estimates a 2.5x to 3.0x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.1x to 2.6x on the GPU's own node)", and the placed row narrows it to one figure each on the next landing.** The labels on its numbers as first written: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound.
The lines the page carries beside it, each labelled:
- The three statements, separate (10.0g item 1): energy resistance (the figures above); economic resistance (the profitability surface of 10.0f item 2, lane 3's first cut, modelled: p* scales as the project cost over the share times the discounted life, and under 5 percent with the per-joule edge; the cheapest attractive project is a USD 20 M DRAM-board design taking the whole chain for three years at about IGN 0.02 to 0.03, at a third 0.055 to 0.10; the SRAM die at N2 0.22 to 0.73; a fixed-lane chip under rotation needs 4x the price of a programmable one; stated as the conditions under which development is attractive); response capability (a passed rotation boundary proves the rotation works, not that hardware dies; the schedule of 10.0d: hourly, weekly, 180-day family, emergency vote; measured per boundary).
@ -471,6 +471,16 @@ The live-state analysis (`livestate.py`, 64 drawn programs, 1,024 waits): under
Two corrections this forces on the served numbers: (1) the honest adversary's base core is the GATED one, k 0.37 at N3 and 0.51 node-for-node, below the 0.56 and 0.78 of 14:0x (those are the GPU-shaped core a maker would not build); (2) placement adds more than the +30 percent estimated at 14:1x: the ungated placed core reads 11.3 pJ against 6.9 synthesised (+64 percent: wires and a 2.5 pJ clock tree). **So until the placed gated rows land (in flight on a rented pod, 17:30 UK) the served figures sit in a bracket, from the synthesised gated rows (4.5 and 6.2 pJ per lane-op: the GDDR7 board 3.3x to 2.9x at the lock at N3, 2.9x to 2.5x node-for-node) to the placed ungated row (11.3 pJ: 2.3x at N3 and 2.0x node-for-node for the base core, the window below it), with the placed gated figure expected near 6 to 8 pJ (approximate): about 2.6x to 3.1x at the lock at N3 and 2.3x to 2.7x node-for-node, the window about 0.3x under the base.** The placed gated rows replace this bracket as the served number when they land, and 10.0h's figures are read from them.
The row to serve at 18:30 UK (the k lane, 17:5x UK; the placement of the gated 64-register core slipped to about 18:30 on a floorplan timing repair, the other five placed rows by 21:00): on the gated 64-register core, synthesis-only, a model never a lower bound (the GDDR7 board at the 5090's 1,300 MHz lock, 2.33 microjoules; E_chip = 0.466 + 102,100 x e_chip; node factors claimed):
| Figure | Chip pJ per lane-op | E_chip microjoules | The edge at the lock | Label |
|---|---|---|---|---|
| node-for-node (N5, ASAP7 x0.70) | 4.3 | 0.905 | 2.6x | synthesised |
| a node ahead (N3) | 3.1 | 0.783 | 3.0x | synthesised, scaling claimed |
| two nodes ahead (N2) | 2.2 | 0.691 | 3.4x | synthesised, scaling claimed |
The placed figure runs 30 to 65 percent over synthesis on this flow (the ungated base came in 64 percent over, wires and a clock tree, which gating removes in part), so the placed gated core is expected at 7 to 9 pJ at ASAP7, which puts the served figures at 2.1x to 2.4x node-for-node and 2.5x to 2.8x a node ahead (approximate until the placed row). **So the served bracket of this section holds and tightens to its lower half, and the honest sentence until the placed row is: "estimates a 2.5x to 3.0x energy-efficiency advantage (2.1x to 2.6x on the GPU's own node)"**, the placed row narrowing it to one figure each. The GPU side of the window is measured (no spill, at most 5 percent per load); the connected-state and multi-family lanes' rows agree with this core within 5 percent.
#### 10.0j Amendment after the landing (16:4x UK): the multi-family adversary lane's first core, and a disagreement between two models that the placed rows settle
The multi-family adversary lane (a1a9876a88f5a72fc; synthesis only, ASAP7 TC, a gate-level random-input VCD; the SRAM macro energy modelled with a band; node factors claimed; for class v7, but it bears on the served window line): one in-order SIMD core with the 64-register window in a FakeRAM 64 x 256 macro per 8 lanes (one 256-bit access serves eight lanes), the imem in two 256 x 34 macros, every unit operand-isolated, a 5-phase single-port slot (throughput bought with lanes, not ports), every bank entry firmware. The genesis-only variant on the class v4 draw: 5.9 pJ per lane-op at ASAP7 (band 5.1 to 7.7), 4.1 at N5, 3.0 at N3; the card pays 10.3 pJ per op on the same draw at the 1,300 lock, so k = 0.40 node-for-node (N5), 0.29 a node ahead (N3), 0.22 / 0.16 at stock. Per family (k N5 / N3 at the lock): the add class 0.45 / 0.33, or 0.36 / 0.26, mul 0.32 / 0.23, mad 0.55 / 0.40, mulhi 0.12 / 0.09, shfl 0.083 / 0.060, the load with the fold 0.43 / 0.31. The whole-hash shadow 0.42 microjoules at N5 (0.31 at N3), so the GDDR7 board reads 2.6x against the 5090 at its lock node-for-node (2.9x a node ahead) and 2.3x / 2.6x against the 5080. **The lane's reading: a re-optimised core sits 15 percent under the k lane's 32-register flop core and 40 percent under its 64-register flop core at the same node, and the register-window knob buys the card nothing once the adversary puts the state in a macro.**
@ -527,6 +537,30 @@ The advantage, separated (the chip at 0.06): the board over a Blackwell owner 0.
**RESPONSE capability:** rotation costs a chip versatility, not life. The 18-family bank costs a chip firmware plus 43 percent of its core cells and 11 percent of its shadow energy, with zero obsolescence credit on any transition in the bank (10.0m); a passed boundary proves the rotation works (10.0d). The window: its form is not the lever (the gated flop file and the macro file agree within 5 percent; the residual over 32 registers 0.3 to 0.4 pJ); its cost to the card is under 1 percent, measured on a 5090 (+0.6 percent) and a 4090 (-0.9 percent) at stock on 8 October (the connected-state lane, replacing "about 0, unmeasured"); and the connected-state class is KILLED (the connected window moves the chip's edge 1.10x and 1.08x against the 1.25x gate on the k lane's re-optimised core; necessity costs a clock-gated file nothing, since it pays per write, not per live register; only the window's width reaches the chip, +0.14 of k).
#### 10.0o Amendment (16:3x UK): the mixed FP32 candidate, KILLED on the GPU budget and the full-board score (`docs/analysis/class-v6/mixed-fp32.md` on class-v6-mixedfp at 255be026; all measured unless marked; for class v7)
The candidate: the class v6 shape unchanged, four FP32 families (fadd, fmul, ffma, fcvt) drawn in the shadow block beside the ten integer families behind `IGNEUM_FG_FP32` (harness only), every result xor-injected, every operand a masked bitcast with the exponent field confined to 96..159 so no input or result is ever denormal, NaN or infinite; round to nearest even, no contraction, no fast-math, written in the emitter for CUDA, OpenCL and Metal; two weights, fp12 (18 percent of the shadow FP on the seed) and fp24 (30 percent).
| Reading | The numbers | Label |
|---|---|---|
| Determinism | bit-identical CPU reference against CUDA on Ada and Blackwell: the pack self-test PASS on three rented cards for every pack, the 2^24 fingerprint equal on the CPU and all three cards for ctrl (5203e444a20bc754), fp12 (d9ddef1fa7a7895a) and fp24 (8fdedbb54ad3614f); Metal and the AMD OpenCL row owed | measured (PROVED on CUDA) |
| Census (sub-version 3, build-4, 256 seeds no era and 256 across eras 0 to 7) | both candidates 256 of 256 both ways, 0 exhausted, r 0.65 to 0.81 against the record's 0.67 to 0.83; the bias instruments fire 7x to 19x the record ((c'') 54 and 88 refusals against 8, (c''') 15 and 19 against 1, the era window-bit test 27 and 35 percent of candidates against 14.5); the F8-form read finds hot items (271 and 740 reads against the control's 29) on 1 of 16 and 2 of 16 seeds: an IEEE result's exponent byte carries 3 to 5 bits of entropy and the xor lands it on address bits 23 to 30 | measured |
| The verifier | the quiet core +5.6 percent (fp12), +8.2 (fp24); loaded, fp24 sits on the 10 ms line | measured |
| The card at stock (the class v5 kit worker, 250 x 2^24, nvidia-smi 1 Hz) | the 5090 3.553 microjoules per hash on ctrl, 4.078 on fp12 (+14.8 percent, 140.7 MH/s held, the card at its 575 W cap), 4.034 capped on fp24 (+13.5 with the clock down 108 MHz); the 4090 4.759, 5.664 (+19.0) and 6.022 (+26.5); v5-genesis +1 percent on both. 52 pJ per FP family op on the 5090, four fifths of it the determinism tax (the four integer ops per operand); the 10 percent budget allows about 11 percent of the shadow FP on the 5090, 9 on the 4090 | measured |
| The chip side (modelled on the k lane's method; its synthesised rows owed) | an FP32 FMA lane on its own is the family a chip undercuts least, k 0.4 to 0.5 at the lock against mad's synthesised 0.20 (the hypothesis's grain of truth), but the drawn op is the FMA plus its masking, which is ARX work on both sides, so the blended k of an ffma family op is 0.19 and an fadd's 0.18: the integer families' own | modelled |
| The full board, E_GPU over E_adversary at the 5090's lock | 3.0x on the record, 3.2x under fp12 (3.2x and 3.4x a node ahead): the candidate raises the chip's edge about 7 percent while costing every card 15 to 26 percent | modelled on measured card rows |
**KILL.** The meaning for v7: FP32 is deterministic across CUDA and the CPU under the stated rules; the determinism tax is the whole economics (any FP form a GPU runs bit-exactly on random registers needs the operand confined, and confinement is integer work at integer k); the exponent-byte bias is a new instrument reading (the F8-form max-item column, 271 and 740 against 29, which no rule reads today) worth a rule in layer 4. The FP32 candidate joins the long program, the select tree, the wide read and the scratchpad in the suite as a negative control (10.0g item 2).
#### 10.0p Amendment (16:38 UK): the 8 GB and 12 GB tiers at the 5.5 GiB floor, the window on them, and proving beside mining (the fleet hand's v6-coexist rows on exact rented cards; the ds55 kit's own worker; measured; logs in the hand's run.log per row)
| Card | The 5.5 GiB dataset (ds55, the pinned class v3 program, 60 s rows) | The 64-register window (the v5 kit worker: hl-reg64c, the full chain; hl-reg64) | Proving beside mining (SP1 on the patched floor, the shard fees-v1-shards2 shard 0, 4.7 M cycles) | Label |
|---|---|---|---|---|
| RTX 3060 12 GB (driver 610, a 170 W limit) | 26.82 MH/s at 117.4 W, 6,129 MiB resident, the 5090's fingerprint 23ced07a4d28b465, PASS: **the 12 GB tier holds at the floor** | hl-reg64c 13.47 MH/s at 120.8 W (87 registers, 16 of 24 blocks per SM, fingerprint MATCH); hl-reg64 13.48 at 119.3 W (104 registers): per load (256 against 128 a hash) 3,448 against 3,433 M loads a second, **the window free per load, +3 percent of watts** | the proof alone 13.2 s VERIFIED at a 7,525 MiB peak and 122 W; together 6,129 + 7,525 = 13,654 MiB against 12,288: TIME-SHARING NEEDED (the live attempt filled the card to 11,893 MiB and the prover died in a device allocation after 34 s while the miner held 26.51 MH/s) | measured |
| RTX 4060 8 GB (driver 570, a 115 W limit; the host's power sensor N/A, so no watts) | 18.84 MH/s, 6,116 MiB resident, the same fingerprint, PASS: **the 8 GB tier holds at the floor** (the question lane 3 left open in 3.3, answered on an exact 8 GB card) | hl-reg64c 9.51 MH/s (87 registers, 20 of 24 blocks); hl-reg64 9.57 (104 registers, 16 of 24): per load 2,435 against 2,412, **the window free per load** | the proof alone 8.2 s VERIFIED at 7,532 MiB; together 13,648 MiB against 8,188: TIME-SHARING NEEDED (the live attempt died in the server's tensor allocation at 7,811 MiB while the miner held 18.83 MH/s) | measured |
What it moves: layer 2's per-tier table (3.3) gains two measured rows at the genesis floor, the 8 GB and 12 GB tiers both holding with the dataset resident (6.1 GB) and the RX 7600 row still owed for the AMD 8 GB case; the window's cost to the card is now measured free per load on Ampere and Ada low tiers as well as on the 5090 and 4090 (10.0e, 10.0n); and the proving statement (10.0f item 4: proving is an opportunity for GPU owners) carries its memory condition: at the 5.5 GiB floor an 8 GB or 12 GB card cannot hold the miner and the SP1 prover at once (13.6 GB together) and time-shares them, while a 16 GB card and up co-resides. Not measured: the core-only beside row (the prove host has no core mode); watts on the 4060. A fault for the floor lane: the served sm_89 tarball (igneum-floor-sm89.tgz) ships the stock SDK server in home/.sp1/bin (the 24 GB gate) while bin/ holds the patched one; the hand copied bin/ over home on the pod, so the 4060 proof rows are on the patched server.
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
| Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |

View file

@ -0,0 +1,28 @@
# The proving payment: the tip-share discrepancy resolved (Igneum 2.0, D4)
8 October 2026, 17:5x UK, the research lane, under the founder's Igneum 2.0 decision (`docs/plans/igneum-2.0.md`, D4: "the economics page's tip-share discrepancy resolved; explicit user-funded proving payment with congestion pricing, burn treated separately; hard cap and no development tax kept"). A design decision on text and spec; the one code change it implies is pinned below and is not made here.
## The discrepancy
The economics page's fee table says the priority fee (the tip) is 80 percent to the block's miner and 20 percent to the developer registrations, which is what the code on Devnet 3 does (`DEVELOPER_SHARE_PERCENT = 20`, the rest to the miner). Spec 5.2 says the 80 percent goes to "the block producer and provers ... in the proportion the proving protocol defines (forward reference)", spec 5.3 speaks of "the provers' part of the 80% tip share", open item O-5.7 leaves that proportion to phase 2, and the page's own "proving-fee market" paragraph says a card's second income includes "the provers' part of the priority fee". So the page contradicts itself and the spec contradicts the code: the provers are promised a share of the tip that no rule sizes and no code pays.
## The decision
1. **The tip stays whole to the block.** The priority fee splits 80 percent to the block producer and 20 percent to the developer registrations, exactly as the code does. No part of the tip reaches the provers. O-5.7 is closed by this decision: the provers' proportion of the tip is zero.
2. **The provers are paid by an explicit, user-funded proving payment with congestion pricing: the proving base fee.** Every transaction already pays `pgas used x f_p`, where `f_p` is the proving base fee that spec 5.1 adjusts per chain block by the EIP-1559 step toward a target of half the proving budget (`B_p / 2`), never below the floor of 5.11. Today that payment is burned. From this decision it is the provers' payment: **90 percent of it is credited to the block's proving pool escrow (`PROVING_POOL_ADDRESS`) and paid out per shard by consensus proving cost under the rules of 5.3; 10 percent is burned.** The price is the congestion price by construction: `f_p` rises when blocks use more than half the proving budget and falls when they use less, so a proving demand spike raises what users pay provers per unit of proving work, which is the signal that brings capacity in (the operator simulation of D4 reads it as its pricing rule).
3. **The burn is treated separately and listed in one place.** Burned: the execution base fee in full (anti-stuffing, unchanged), 10 percent of the proving payment, the unregistered developer share of the tip, and 10 percent of IGN-settled external jobs (5.4). Nothing else.
4. **The hard cap and the absence of a development tax are untouched.** No new emission, no change to the 20 percent proving-pool share of the subsidy (2.5), no address that any team controls.
Why the 10 percent burn on the proving payment: the burn of the proving base fee was the rule that made wash pgas a guaranteed loss (5.1). With 90 percent of it routed to the block's provers, a miner who is also a prover of its own block could recover part of a stuffed block's proving payment; sortition (eight eligible provers per shard, 5.3) makes that recovery a share of the pool's weight at best, and the 10 percent burn plus the execution base fee burned in full keep stuffing a loss at every weight. The 90 and 10 mirror the external job split of 5.4, so a prover's two incomes carry one rule.
## What a prover earns, in one line
Per block: its shards' part of 20 percent of the block subsidy (2.5, 5.3) plus its shards' part of 90 percent of the block's proving payment (`pgas used x f_p`); per job: 90 percent of the external job fee (5.4). Nothing from the tip.
## The code pin (not made here)
`igneum/exec/src/executor.rs` (the proving base fee debit at 357 and 360 on Devnet 3): credit 90 percent of `pgas used x f_p` to `PROVING_POOL_ADDRESS` and debit the remaining 10 percent to no one, instead of debiting the whole to no one; the pool's per-shard payout (5.3) then carries it with no further change. Behind an activation constant on the devnet objects like `proving_v1_activation_daa`. Until it lands the economics page says "designed, not in the code" on that row, as it does for the external job split.
## Where the text changes
Spec 5.1 (the proving-cost gas row and the burn sentence), 5.2 (the 80 percent recipient and the forward reference), 5.3 (the provers' income), 06 O-5.7 (closed); the economics page's "Where fees go" table (the base-fee row split into the two dimensions, a proving-payment row, the duplicated tip row removed) and its "proving-fee market" paragraph.

View file

@ -13,7 +13,7 @@ Designed. Every transaction pays a base fee in both gas dimensions:
| Execution gas | EVM execution, Ethereum's rule | Ethereum's EIP-1559-style base fee over the ordered sequence |
| Proving-cost gas | Proving cycles the transaction will cost the provers | A second base fee `f_p`, adjusted per chain block by the same EIP-1559 step as `f_e`: toward a target of `B_p / 2` of proving gas used, denominator 8, never below the floor of section 5.11 (one definition, 5 October 2026, ledger P14; `next_base_fee` in `igneum/exec/src/executor.rs`, applied per chain block in `service.rs`). The unproven backlog does not move `f_p`; it halves `B_p` (design 4.3, the backlog rule), which raises `f_p` through the step. No smoothing over the difficulty window is implemented or specified: the two-dimension step is per chain block |
The base fee in both dimensions is **burned in full**. A miner cannot stuff blocks with its own transactions for free; wash gas loses its whole base fee (ledger E3). The proving-cost budget per block is a consensus constant set from measured prover throughput (phase 2 gate: one shard on a 12 GB card in about 20 s, Target, unmeasured, ledger P1), so a transaction that is cheap to run and brutal to prove cannot stall the provers for everyone. How the node folds the proving-cost dimension into the quoted gas price so `eth_estimateGas` keeps working is fixed in section 7.1 (ledger P5, closed 3 October 2026).
The execution base fee is **burned in full**; the proving base fee is the provers' payment from 8 October 2026 (90% to the block's proving pool, 10% burned; `docs/design/proving-payment.md`, Igneum 2.0 D4), which keeps the anti-stuffing property below through the burn and the pool's sortition. A miner cannot stuff blocks with its own transactions for free; wash gas loses its whole base fee (ledger E3). The proving-cost budget per block is a consensus constant set from measured prover throughput (phase 2 gate: one shard on a 12 GB card in about 20 s, Target, unmeasured, ledger P1), so a transaction that is cheap to run and brutal to prove cannot stall the provers for everyone. How the node folds the proving-cost dimension into the quoted gas price so `eth_estimateGas` keeps working is fixed in section 7.1 (ledger P5, closed 3 October 2026).
## 5.2 Priority fee split
@ -21,7 +21,7 @@ Designed. The priority fee (tip) of every executed transaction splits:
| Share | Recipient | Rule |
|---|---|---|
| 80% | Block producer and provers | The producer of the block in which the first copy executed (section 2.6) and the provers of that block, in the proportion the proving protocol defines (forward reference) |
| 80% | Block producer | The producer of the block in which the first copy executed (section 2.6). No part of the tip reaches the provers: O-5.7 is closed at zero (8 October 2026, `docs/design/proving-payment.md`); the provers' user-funded payment is the proving base fee (5.1, 5.3) |
| 20% | Developer | Attributed per call frame by gas consumed, to the developer address registered for the called contract at deployment. A frame in an unregistered contract sends its share to the burn |
No part of the tip is burned by rule; the burn is the base fee (5.1) and the unregistered developer share.
@ -34,7 +34,7 @@ Self-dealing: a developer who also mines the including block collects 80% plus 2
## 5.3 Proving pool
Designed. The 20% emission share (section 2.5) and the provers' part of the 80% tip share are paid per block as a fixed amount for that block, divided among the block's shards by consensus proving cost, so a stuffed block earns no more than an honest one. Shards are not claimed first-come and carry no bond: each shard is assigned by sortition to 8 eligible provers for a 10-s exclusive window, then open to anyone, and the first valid proof included in a block is paid (section 7.2, decided 3 October 2026, ledger P8, C9). The parameters 8 and 10 s are set on the phase 4 devnet (O-5.1). A withheld shard costs nothing to bond against because nothing waits on an assigned prover: an unproven block delays only its proof; execution and the 30-s lock do not wait for it (ledger P9). The bond, slashed on a bad or late proof, remains in the external job market (5.4), where a customer does wait; its size and timeout are Open (O-5.6).
Designed. The 20% emission share (section 2.5) and 90% of the block's proving payment (`pgas used x f_p`, the congestion-priced proving base fee of 5.1; 8 October 2026, `docs/design/proving-payment.md`, the code pinned) are paid per block, the emission share as a fixed amount for that block and the proving payment as the block's own, divided among the block's shards by consensus proving cost, so a stuffed block earns no more than an honest one. Shards are not claimed first-come and carry no bond: each shard is assigned by sortition to 8 eligible provers for a 10-s exclusive window, then open to anyone, and the first valid proof included in a block is paid (section 7.2, decided 3 October 2026, ledger P8, C9). The parameters 8 and 10 s are set on the phase 4 devnet (O-5.1). A withheld shard costs nothing to bond against because nothing waits on an assigned prover: an unproven block delays only its proof; execution and the 30-s lock do not wait for it (ledger P9). The bond, slashed on a bad or late proof, remains in the external job market (5.4), where a customer does wait; its size and timeout are Open (O-5.6).
Proving v1 (section 7.8, 5 October 2026, Implemented behind `proving_v1_activation_daa`): from the switch, `proving_v1_aggregator_share_bps` of a block's fixed amount (a tenth, the founder's decision at 0.3.11) goes to the aggregator whose segment record attests the block, the rest to the shards as before; a block in a segment that stays unproven past `proving_v1_unproven_daa` pays no aggregator share.

View file

@ -93,7 +93,7 @@ An item closes when its measurement is in `docs/bench-log.md` or its decision is
| O-5.4 | Every genesis contract's upgrade and key policy (ledger G4); the development fund contract is gone (fund removed 3 October 2026) | Publish before launch | 4 |
| O-5.5 | The proving-cost gas dimension: the metering table per opcode and precompile, and the per-block budget from the phase 2 measurement | Phase 2 benchmark; execution-layer specification | phase 2 |
| O-5.6 | External job bond size and claim timeout (ledger P9); shards carry no bond since the sortition rule of section 7.2 | Set on the phase 4 devnet | 4 |
| O-5.7 | The provers' proportion of the 80% tip share (section 5.2) | Defined by the chunked proving protocol | phase 2 |
| O-5.7 | The provers' proportion of the 80% tip share (section 5.2) | Closed 8 October 2026: zero; the provers are paid by the proving base fee, 90% to the block's pool and 10% burned (`docs/design/proving-payment.md`, Igneum 2.0 D4) | decided |
| O-5.8 | Developer registration format at deployment and the re-registration transaction (section 5.2) | Execution-layer specification | decision, execution engineer |
| O-5.9 | Mining against proving under shocks: external proving pays 10x, the IGN price falls, a large operator leaves, assignments go unfulfilled, clients maximise profit (ledger E12); design R8 covers the fee switch only and has not run | R8 simulation extended with the five shocks, then the phase 4 devnet with profit-only prover clients: backlog depth, time to clear, `f_p` and `B_p` paths, income per card. Sweep 5 October 2026: the simulation half ran in `sim/economy` (scenarios b, d, e: external 10x with a 70% price fall, the 20% operator leaving, a 30% operator never fulfilling; no backlog, every block proven within 60 s, hash trough 82% and 75%); the devnet half with profit-only clients is fud-fixes row 126 | 4 |
| O-5.10 | No single figure shows every payment route (emission, base fee, priority fee, external jobs at launch and after the bridge, the client dev fee) with currency, recipient, fee and burn (ledger E13) | Draw it, one route per row, in the litepaper Economics section and the customer brief; operator revenue never summed with protocol revenue | decision, execution engineer; before public repo |