Merge counter-asic-4 a39ab8ec into master (gate: green on a39ab8ec, recorded by tools/ci/pre-push.sh; landed on the build mirror)

This commit is contained in:
igneum-labs 2026-10-08 15:25:52 +00:00
commit b4de1eccfd

View file

@ -439,6 +439,8 @@ The conditions, read off the surface, which are the economic-resistance statemen
#### 10.0h The served text (for the site audit lane, served as written; each number's label; what is withdrawn)
**The claim statement above the sentence (the third review, 10.0l item 3), served verbatim:** "Igneum remains competitive on accessible commodity GPUs even when specialised mining hardware is assumed to exist, remain compatible and seek profit; its security does not rely on identifying that hardware or retiring it through emergency changes." Under it the three statements (energy, economic, response capability), rotation named as an optional improvement to the baseline, not the mechanism; the economic statement reads from the coexistence model (10.0l item 2, by 21:00 UK) once it exists, and until then names the model as owed rather than any threshold.
**The sentence, as the external review words it (10.0f item 5), served verbatim, with one word made honest (the site audit lane's read, 16:5x UK: class v4 and v5 have eight registers per lane and the window is class v6's new core shape, so "retains" is read as "retains across every rotation"):** "Class v6 adopts the 64-register window and retains it across every rotation. Current modelling estimates a 2.2x to 2.4x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.0x on the GPU's own node). The long-program and select-tree proposals were rejected. Economic resistance depends on development cost, deployment economics and productive hardware lifetime; family transitions receive an obsolescence benefit only where a loss of competitiveness is demonstrated; programmable multi-epoch designs are included in the assessment."
**Where the figures come from (the coordinator, 16:1x UK): the served sentence's numbers are read from the k lane's PLACED GATED rows (17:30 UK; 10.0i), not from 725d2945's close and not from the 16:0x synthesis; until they land the figures sit in 10.0i's bracket, and the serve slips to 18:30 if the rows are late.** The labels on its numbers as first written: "2.2x to 2.4x" is modelled (the GPU side measured: the RTX 5080 at its 1,100 MHz lock 2.06 microjoules per hash and the RTX 5090 at its 1,300 MHz lock 2.33, both on class v4, PC 1 and rented pods, 8 October 2026; the chip side the k lane's synthesised 8-lane sequencer core with the 64-register window on ASAP7, scaled to N3 on TSMC's headline factors, claimed; the chip's memory the chip model's GDDR7 board, modelled; the card's cost of the window measured at stock on a rented 5090 and 4090 at 16:4x UK, within 5 percent per load with the liveness chain, no spill). "2.0x on the GPU's own node" is modelled (the same core node-for-node, k 1.09). Against the 32-lane window core the adversary would build the same figures read about 2.4x to 2.6x a node ahead and 2.0x to 2.2x node-for-node (synthesised, pending the re-optimised row by 18:00 UK); the served sentence's range is kept as the review wrote it and the 32-lane rows sit beside it on the page as the pending row. The window's k is synthesis-derived and not a lower bound.
@ -451,7 +453,7 @@ The lines the page carries beside it, each labelled:
- The evaluation rule (10.0g item 2) with the harness and scoring-rules links: the minimum over workloads of the maximum over free adversarial designs of E_GPU over E_adversary, under the 10 percent GPU-cost budget at the lock, the verifier limit, cross-vendor correctness and hardware accessibility; the rejected long program, select tree, wide read and scratchpad as negative controls with their measured rows.
- The next programme (10.0g item 7): the connected-state experiment, the mixed integer and FP32 candidate, the multi-family programmable adversary; rotation is not expected to deliver the missing joule.
Withdrawn, to be struck everywhere it appears: the lifetime claim ("under 1x per hash over its 180-day life"; 10.0f item 1); the USD 300 M and USD 340 M lines and any "the chain stays below X" wording (10.0f item 2); any chip-arrival probability (10.0g item 5); "Monero's hash never changes" (10.0g item 6); the 725d2945 sentence ("under 3x per joule at the knee, under 1x per hash over its life") and this lane's section 9 proposal; W = 8 as a next width (W = 4 at genesis, 8 measured not free, 16 never).
Withdrawn, to be struck everywhere it appears: any "capex wall" or threshold as the economic headline (the coexistence model is the argument; the USD 23 M and 340 M rows are its rows, 10.0l); the lifetime claim ("under 1x per hash over its 180-day life"; 10.0f item 1); the USD 300 M and USD 340 M lines and any "the chain stays below X" wording (10.0f item 2); any chip-arrival probability (10.0g item 5); "Monero's hash never changes" (10.0g item 6); the 725d2945 sentence ("under 3x per joule at the knee, under 1x per hash over its life") and this lane's section 9 proposal; W = 8 as a next width (W = 4 at genesis, 8 measured not free, 16 never).
#### 10.0i The adversary's re-optimised core, and the bracket the served figures sit in (the k lane, 16:0x UK; the coordinator's order 16:1x: the placed gated rows at 17:30 UK are the figures to serve, not 725d2945's and not this synthesis; the 18:00 serve slips to 18:30 if they are late)
@ -469,6 +471,37 @@ The live-state analysis (`livestate.py`, 64 drawn programs, 1,024 waits): under
Two corrections this forces on the served numbers: (1) the honest adversary's base core is the GATED one, k 0.37 at N3 and 0.51 node-for-node, below the 0.56 and 0.78 of 14:0x (those are the GPU-shaped core a maker would not build); (2) placement adds more than the +30 percent estimated at 14:1x: the ungated placed core reads 11.3 pJ against 6.9 synthesised (+64 percent: wires and a 2.5 pJ clock tree). **So until the placed gated rows land (in flight on a rented pod, 17:30 UK) the served figures sit in a bracket, from the synthesised gated rows (4.5 and 6.2 pJ per lane-op: the GDDR7 board 3.3x to 2.9x at the lock at N3, 2.9x to 2.5x node-for-node) to the placed ungated row (11.3 pJ: 2.3x at N3 and 2.0x node-for-node for the base core, the window below it), with the placed gated figure expected near 6 to 8 pJ (approximate): about 2.6x to 3.1x at the lock at N3 and 2.3x to 2.7x node-for-node, the window about 0.3x under the base.** The placed gated rows replace this bracket as the served number when they land, and 10.0h's figures are read from them.
#### 10.0j Amendment after the landing (16:4x UK): the multi-family adversary lane's first core, and a disagreement between two models that the placed rows settle
The multi-family adversary lane (a1a9876a88f5a72fc; synthesis only, ASAP7 TC, a gate-level random-input VCD; the SRAM macro energy modelled with a band; node factors claimed; for class v7, but it bears on the served window line): one in-order SIMD core with the 64-register window in a FakeRAM 64 x 256 macro per 8 lanes (one 256-bit access serves eight lanes), the imem in two 256 x 34 macros, every unit operand-isolated, a 5-phase single-port slot (throughput bought with lanes, not ports), every bank entry firmware. The genesis-only variant on the class v4 draw: 5.9 pJ per lane-op at ASAP7 (band 5.1 to 7.7), 4.1 at N5, 3.0 at N3; the card pays 10.3 pJ per op on the same draw at the 1,300 lock, so k = 0.40 node-for-node (N5), 0.29 a node ahead (N3), 0.22 / 0.16 at stock. Per family (k N5 / N3 at the lock): the add class 0.45 / 0.33, or 0.36 / 0.26, mul 0.32 / 0.23, mad 0.55 / 0.40, mulhi 0.12 / 0.09, shfl 0.083 / 0.060, the load with the fold 0.43 / 0.31. The whole-hash shadow 0.42 microjoules at N5 (0.31 at N3), so the GDDR7 board reads 2.6x against the 5090 at its lock node-for-node (2.9x a node ahead) and 2.3x / 2.6x against the 5080. **The lane's reading: a re-optimised core sits 15 percent under the k lane's 32-register flop core and 40 percent under its 64-register flop core at the same node, and the register-window knob buys the card nothing once the adversary puts the state in a macro.**
The disagreement, carried as two models until the placed rows settle it: the k lane's SRAM-banked time-multiplexed window (10.0i) is modelled at 8.5 to 10.5 pJ per lane-op, NOT cheaper than its gated flop window (6.2), with the RTL form in synthesis on a pod as a check; the multi-family lane's macro window reads 5.9 (band 5.1 to 7.7) for the whole core. The two differ in the macro's energy model and the access pattern (one 256-bit access per eight lanes against a per-lane port), and in whether the window's whole state is live across every wait (10.0i: 61 of 64 values necessary, so no lane can skip the access). Until the k lane's placed gated rows (17:30 UK) and the multi-family lane's placed 8-lane core (18:30 UK) land, the served window line is read with this caveat: **the window's measured cost to the card is under 5 percent and its defence against a flop-file adversary is about 0.13 of k; against a macro-file adversary it may be close to nothing, which is why the next programme's multi-family adversary is the lane that decides whether the window stays a defence or only a cost.** The 18-family variant (+43 percent of cells on the same microarchitecture), the live-family mixes and the 32-lane cores follow by 18:30 with the board and lifetime rows.
#### 10.0k Amendment (16:56 UK): the connected-state lane's first rows (for class v7; branch class-v6-connected, `docs/analysis/class-v6/connected-state.md` by 17:45 UK)
The variant cs64s27x16 (a research class): a 64-register window per lane; per step one 4-byte load from a window register, then 27 distinct instructions run 16 passes, the next address written by the block's last instruction from a fresh spine (add, sub, xor, shfl only), lossy ops feeding the next injection, every register covered by an injecting write, the result folding all 64 registers; 128 loads and 55,296 ALU per hash (class v5 128 and 55,680), text 448 per iteration (v5 320); no long program, select tree, wide read or scratchpad.
| Reading | The numbers | Label |
|---|---|---|
| Liveness (seed igneum-v6c/0) | 63 of 64 registers necessary for a later address at every address until the last iteration's tail; the result reads 64; an address depends on 1 to 15 registers (mean 11.2) of the previous address point; a block touches 16 to 24 registers; reads per register 384 to 3,208 per hash | measured |
| What that means | the window is live state a specialist must hold, but the per-step chain is narrow (about 11 of 64), so the obvious re-optimisation is a two-level file, about 20 hot per step | the lane's reading |
| Census (the sub-version 3 harness, (c''') on) | no era 256 of 256, 0 exhausted, 0.28 rejections per candidate, mean attempt 0.39 (the control mx8+sh256x27: 0.67, 2.0); eras 0 to 7 with the window-bit refusal 256 of 256, 0 exhausted, mean attempt 1.7 (the control 4.9), 306 window-bit refusals (the control 218); the F8 form at 16 x 2^20 0.999 to 1.002 of uniform (the control the same) | measured |
| What that means | it censuses at least as well as v5; it carries v5's product-bit class, which layer 1's index fold removes for both | the lane's reading |
| The card at stock (a rented 5090 at the 575 W cap, the class v5 nvcc harness, 250 x 2^24, the same card for both) | cs64 64.93 MH/s at 574.8 W against v5-genesis 65.30 at 574.8 W: +0.6 percent of energy per hash for the window; ptxas 80 registers per thread (v5 48), 0 spills, the same occupancy | measured |
| What that means | the window is nearly free on the card at stock (the budget 10 percent); the 4090 row by 17:15, the PC 1 lock row owed (PC 1 booked to 17:40); the chip score waits on the k lane's re-optimised pJ and k; KEEP or KILL at 18:00 | owed |
For the window line of the close: this is a second measured card row (after 10.0e's) that the window costs the card under 1 to 5 percent, and a second liveness read (63 of 64) that a specialist must hold the state; the narrow per-step chain (about 11 of 64) is the first sign of the re-optimisation 10.0i and 10.0j are arguing over, a two-level file, and the chip score at 18:00 is what the window's defence is worth against it.
#### 10.0l The third external review, taken (the coordinator, 17:0x UK; the founder accepts it; it reframes the objective and overrides 10.0f to 10.0h where they differ)
**The objective is durable GPU competitiveness, not chip destruction.**
**(1) The definition.** "Ordinary GPUs remain competitive" is measured on a published reference population (several vendors, memory sizes and generations, including used cards: the discrete-GPU cohort of 10.0g item 3, Apple reported and not headlined) in two situations: the existing owner (operate, after electricity, wear, fees and alternatives) and the new entrant (buy, operate, resell). The central measure is cost per ACCEPTED unit of work: annualised hardware plus power plus hosting, failures and fees, over accepted work (including rejects, downtime, epoch preparation and propagation). The hardware advantage is kept separate from the operating advantage (a 1.5x chip at USD 0.06 per kWh against a GPU at 0.25 is 6.25x on electricity alone). The target: **"across declared hardware and operating-cost conditions, accessible GPUs retain competitive economics and specialised hardware does not create an overwhelming additional barrier to entry".** What this document already holds toward it: the per-tier joules (10.4, measured on 14 card classes at stock and at the knee where the lock is allowed), the epoch costs (10.0d: the re-tune seconds and watts, the verifier ms per epoch build), the card's measured costs of the window and the floor (10.0e, 3.3); what it does not yet hold: the accepted-work denominator (rejects, downtime, propagation) per tier, the hosting and failure terms, the used-card prices and resale, and the electricity axis, which the coexistence model takes.
**(2) The coexistence model** replaces the capex wall as the economic argument: a new document, `docs/analysis/class-v6/coexistence-model.md`, by 21:00 UK, with its first run on today's measured rows, labelled modelled. A five-year model with the development cost SUNK as a mandatory case (the opponent covers manufacturing, deployment and operation only); private mining and hardware sales; several productive lifetimes; growing and shrinking networks; falling token revenue; changing difficulty; used-GPU entry and exit; GPU generation upgrades with purchase and resale on both sides; cheaper derivative chip designs; and miners allowed to react. Its outputs: the cost advantage, the replacement economics, the accessible supply, the break-even electricity prices, and the network's dependence on individual suppliers. Success: a specialised supplier earns a normal return and GPUs stay close enough in total cost, obtainable and useful outside mining that entrants still compete. Failure: a supplier operating privately at much lower cost exhausts competitors' margins. **The USD 23 M and 340 M rows of 10.3, and lane 3's surface of 10.0f, become rows of that model and are never the headline.**
**(3) The claim statement:** "Igneum remains competitive on accessible commodity GPUs even when specialised mining hardware is assumed to exist, remain compatible and seek profit; its security does not rely on identifying that hardware or retiring it through emergency changes." Under it the three statements of 10.0g item 1 (energy, economic, response capability), with rotation named as an optional improvement to the baseline, not the mechanism: 10.0d's schedule is what the rotation does when it runs, and the claim stands without it.
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
| Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |