igneum/docs/analysis/class-v6/invention.md

65 KiB

Class v6 invention lane: layers beyond the four, each priced against the chip that stores the dataset

8 October 2026, first cut 11:1x to 13:xx UK (BST, the Mac's clock; the boxes print CEST, one hour ahead, and nothing below is stated in box time), branch class-v6-invention from the mirror's master at 17d61d52, research lane C under the Counter ASIC coordinator. The founder's word of 11:1x UK: class v6 is declared with four layers as its spine (per-era draws of the released parameters; the state-derived dataset's size tracking chain state with a floor; scheduled family epochs every 180 days with no release; the acceptance floor and the F8-form uniformity test generalised to each era's draw), and deep research opens: "see if anything can be optimised, added or invented". This file is the invention lane's answer: every candidate layer beyond the four as one paragraph, one known-failed test stated as a harness run, and one chip-model row, then the ranked list of what v6 should add. The synthesis is the research lane's docs/design/class-v6-rotating-family.md (branch counter-asic-4), which reads this file.

Status: a research document and a gate plan. No consensus code this week; nothing here touches the devnet, the testnet object, the spec or any served page. Every number carries a label: measured (a run on a named box or card, the log named), modelled (arithmetic on the chip model's cited figures, docs/analysis/chip-model-v3.md section 5), claimed (a vendor's or an author's figure, URL and date), approximate (from memory or a scaling). The rule every candidate is held to: reject what costs GPUs more than it costs chips, and say so with the number; keep what raises a chip's k or capex or shortens its useful life more than it raises every GPU tier's cost.

0. One page

The frame is last night's identity (docs/analysis/counter-asic-4-research.md section 2): against the chip anyone builds, a GPU's memory system without the GPU (chip-model-v3 section 5.5, the f = 1 chip), the per-joule edge is (E_card + F) / (E_mem + k F), with F the premium per hash the card pays for forced work and k the chip core's energy per forced op over the card's. Zero premium is zero forcing. No hash-side design reaches a chip edge of about 2x at zero premium; the premium-free floor is the card's own whole-card energy over the chip's memory energy, 3.6x on a 5090 at its knee (measured card, modelled chip), and the class v4 shadow buys 2.1x at k = 1 for 82 to 90 W on a 5090. The four layers of v6 render a fixed-function chip useless on its release day and move nothing in the identity for the GPU-like chip. So an invented layer can do one of four things, and only four: raise k (nothing on measured rows reads k above 1, section 15.1a of the research file); raise the chip's capex or project cost (the one lever the identity does not see); shorten the chip's useful life (layers 1 and 3 already do this for fixed silicon); or cost the honest cards less at the same F. Every candidate below was read against those four.

What was measured today (all on igneum-build-1, the counter-asic-4 crate at 5984ffab, logs under /srv/builds/v6-invention/):

Reading Number Label
The sound per-load shadow (one pass of a 256-instruction sub-block after every load, mx8+shl4096x1), the form 20.2a-close named and never drew 234 of 256 seeds accepted within the 32-attempt cap, 0.927 rejection per candidate, mean accepted attempt 9.7; the iterated 16 x 27 form on the same seeds 76 of 256, 0.989 per candidate measured (section 3.1)
The same at class v4's instruction count (mx8+shl2304x3: 16 sub-blocks of 144, three passes, 55,296 shadow instructions per hash) 224 of 256 seeds, 0.935 per candidate measured
The verifier on the sound forms, cold warp on core 40 with core 88 loaded (the ladder's method) shl4096x1 8.28 to 8.55 ms; shl2304x3 8.44 to 8.82; class v4's shape 8.33 to 8.63 on the same core in the same minutes; all under the 10 ms gate measured (section 3.2)
The F8-form uniformity read at 2^20 nonces on 64 seeds (v6inv-uniform, the crate's own trace_load_indices in the class's execution order) the sound forms read clean: shl4096x1 0 of 60 seeds over 1.2x at the top 0.1 percent of items (max 1.034x, one item at 251 reads against the control's 29 on seed 23), 0 sites under the (c''') floor of 0.995 (min 0.99707); shl2304x3 0 of 61 over 1.02x (max 1.0019), min site 0.99962; the raw class v4 shape drawn through the same v2/v3 rule (no (a'), no (c''): the pre-amendment class) 41 of 64 seeds over 1.2x at 7x to 36x the control, items at 350,000 to 396,564 reads, min site ratio 0.093: AP-F8-1's hot-set class, the harness's known-failed firing on real programs measured (section 3.5)
The per-load prototype's acceptance under a drawn era 0 of 256 seeds on every per-load form: 5,536 of 7,862 bias rejections name index bit 26 or 27 "set in 0 or 16,384 of 16,384", the era WINDOW's fixed top bits, an instrument fault; so 20.2a-close's "1.4 percent accepted, 42 of 64 seeds exhaust" (run across drawn eras) measured the instrument, not the construction measured; a correction for the research file (section 3.3)
The two packs on the RTX 5090 (PC 1, the hash lane's job run-ca4-pc1-v6-packs-5090-20261008, 12:14 to 12:20 UK, self-test PASS and the pinned fingerprint on every row) at class v4's own instruction count the sound per-load form (shl2304x3, 55,296 ops) costs the card 19 W MORE than the whole block unlocked (483.6 against 464.6 W at 135.85 against 137.62 MH/s) and 6.6 W more at the 1,300 MHz lock (303.4 against 296.8 W); per shadow instruction 19.7 pJ unlocked against the whole block's 17.2, the one-pass form (shl4096x1, 32,768 ops) 20.8: the per-load placement pays 15 to 20 percent more per instruction on this card, the OPPOSITE sign from the 16 x 27 v2 export's 13 to 14 W under measured (section 3.4); the GPU column of rank 1 corrected

The ranked list of what v6 should add beyond the four (section 4 has every column):

Rank Layer What it does to the chip What it costs the cards (5090 / 5070 Ti / M5 Max) Status
1 Layer 5: the shadow placed per load, one pass of a long sub-block (the capex lever: the chip's core must sit inside every read's dependency, so controller, lanes and PHY share one N5-class die or an interposer) energy edge unchanged (k is k whichever die the core sits on); project cost about USD 30 M to about 60 M, the break-even market cap about USD 100 M to about 200 M (the mission lane's model, modelled); capex per MH/s 3.0 to 4.3 USD (modelled) 5090: the sound form at class v4's instruction count measured +19 W (4 percent) unlocked and +6.6 W (2 percent) at the 1,300 lock against the whole block, rate -1.3 percent, 15 to 20 percent more per shadow instruction (PC 1, 12:1x UK) / no row, the 4070's 30 W premium at its tune point as the proxy / the 16 x 27 form -0.5 percent of rate measured; the sound form's Metal footprint OWED acceptance measured today (0.927 per candidate, P(exhaust at 256) about 4 x 10^-9); the sub-version 3 dataflow rule in execution order is the fix that brings it toward v4's 0.68 (gate plan, section 2.1)
2 Layer 6: register-file width drawn per era (8 to 32 registers per lane), the link tax on layer 5 with layer 5, the lane state crossing the controller twice per read grows from 64 B to 128 to 256 B: 2.2 to 9 TB/s of die-to-die traffic at the 5090's read rate, past any one-stack interposer, which closes the "or an interposer" branch and forces the single N5 die (modelled); alone, nothing 0 rate on every card while latency-bound (a GPU lane holds up to 255 registers; 7,262 lanes x 256 B is 1.9 MB against the 5090's 43 MB of register file, approximate); the verifier's register-major arrays 4x (unmeasured, under 0.1 ms by the op law); no pack form today (the register count is a generator constant) modelled; the pack form is a generator change with its own census (gate plan)
3 Layer 7: warp-uniform data-dependent block selection (which of B shadow sub-blocks runs next is chosen by a warp-reduced register value, uniform across the 32 lanes, so no divergence) nothing for the GPU-like chip (a sequencer already); an FPGA overlay or a fixed pipeline must hold all B blocks for one block's throughput: B x the shadow's LUT area (approximate); shortens a per-epoch bitstream's worth one uniform indirect branch per iteration: about 0 (the sel register already does this for the immediates); verifier 0 modelled; a generator change behind a pack (gate plan); rank 3 because it moves the FPGA lane only
4 The reserve ordered by hardware orthogonality (layer 3 as it stands, with the order fixed: shuffle-crossbar families, then byte-permute, then popcount and priority encoder, the int8 tile last) a chip pre-wires every family for about USD 4 of N5 (modelled); the order makes the first unlocks the ones a 12-op datapath lacks most 0 at 4 points (measured step costs under 1 percent of ALU time on every vendor) an ordering rule inside layer 3, not a new layer; the history's addition 6

Rejected with the number, each in section 2: per-lane data-dependent branches (divergence costs the card, a chip nothing); reads tied to the shard proof per block (a refresh per block is 0.3 to 1 W on a chip, 1.3 percent of a 4090's hash time); randomised memory topology (a chip's address decoder permutes its lines for nothing; the stride and interleave are already drawn); the VRAM-size ratchet as a lever (a chip buys 24 to 32 GB that the 8, 12 and 16 GB tiers cannot: it retires cards first); proof-carrying hashes sampled by the pool (the chip holds everything the witness proves); prover-gated eligibility (proving is 1.6 kW network-wide at any hash rate, 0.7 percent of the hash's energy at 100 GH/s); time-locked parameter commitments beyond the era VDF (the drawn band is firmware; the 2-hour lead already denies the fixed chip 180 days); a fraction of reads derived from the cache (the 5090 loses about 30 percent of rate, the chip 7 percent of energy); row-straddling reads (the w64 regime, 47 percent of the 5090's rate); a refresh per block (dead by arithmetic, the research file's row 7).

Per tier, in one line each: a home miner on any card sees no change from anything here today (nothing ships; the devnet pays nothing); the 5090 tier's one number is that the sound per-load placement costs it 4 percent more watts than class v4's whole block at the same work unlocked and 2 percent at the knee (19 and 6.6 W, measured 12:1x UK; the dead 16 x 27 form had read the other way); the 5070 Ti has no measured row in this lane (the 4070's rows are the proxy, approximate); the Apple tier's open question is the inline footprint of a 4,096-line block per iteration (the 1,024-line block cost the M5 Max 17 percent on 6 October, measured), which decides whether rank 1 needs a block-shape cap for Apple; a pool user sees nothing; a node verifies the per-load forms in the same 8.3 to 8.8 ms as class v4 (measured); a chip maker sees its project forced onto one advanced die by rank 1 and its interposer escape closed by rank 2.

1. The frame, and what a candidate must do

Term Value, 5090 at the 1,300 MHz knee Label Source
E_card, class v3 1.67 microjoules (127.3 MH/s at 213.0 W; 134.6 at 223.3 on the efficiency pass) measured research file 20.3 and the status file's efficiency pass
E_mem, the f = 1 GDDR7 chip 0.466 microjoules (16 devices, 166 MH/s at 78 W); one HBM3 stack 0.321 modelled chip-model-v3 5.4
F, the class v4 shadow's premium 0.652 microjoules (82.8 W over 126.9 M hashes x 102,100 ops: 6.4 pJ per counted op) measured research file 20.3
The edge at zero premium 3.6x (GDDR7), 5.2x (one HBM3 stack) measured card, modelled chip section 2 of the research file
The edge with the shadow 2.1x at k = 1, 2.9x at k = 0.5, 3.5x at k = 0.3; the honest k band for an ALU-shaped core 0.3 to 0.8 modelled on measured rows 20.4
The chip's capex USD 470 of memory, controller and board per 166 MH/s: 2.8 USD per MH/s; plus the shadow core USD 25 to 40 (3.0 to 3.1); plus an interposer USD 200 (4.3); a 5090 at MSRP 14.7, at the 2026 street price about 29 modelled; the card price measured 16.1
The project cost and the break-even cap a 28 nm controller USD 5 M (cap about 17 M); plus an N5 shadow core USD 30 M (cap about 100 M); the core forced onto the controller's die or an interposer USD 60 M (cap about 200 M) modelled (the mission lane's model, s = 0.30) 16.2

The four doors, and the one each candidate must walk through:

  1. Raise k. Measured on the 5090 (15.1a): the int32 ALU op 6.2 to 11.3 pJ against a 5 nm SIMD array's 2 to 5 (k 0.3 to 0.8); the shuffle 29 to 56 pJ against a crossbar's about 20 (k 0.4 to 0.7, and the card pays 5x the add per op); the int8 tile 0.8 to 4 pJ per MAC against an array's 0.04 to 0.4 (k 0.03 to 0.3); an L2 hit 1.4 to 2.4 nJ against on-die SRAM's 0.2 to 0.5 (k 0.1 to 0.3). Nothing reads above 1. A candidate that claims to raise k must name the block and the measured GPU cost per op it rests on.
  2. Raise the capex or the project. The hash's dataset is a card's worth of DRAM and its work a fraction of a card's logic, so per-unit capex cannot pass about 4.3 USD per MH/s (16.1); the project cost is the lever that moves the break-even cap, and the only mechanism found for it is forcing the shadow core into every read's dependency (16.2). A candidate here is priced by which die it forces.
  3. Shorten the useful life. Layers 1 and 3 kill fixed silicon at the first draw outside its wired value; a GPU-like chip's life is its memory's and its node's. A candidate here must move the GPU-like chip, or it is layer 1 again.
  4. Lower the honest card's cost at the same F. The operating point (the knee lock) is the miner's lever, not the protocol's; a protocol lever here is a shape that runs cheaper per instruction on the card (the 16-instruction block effect, 13 to 14 W measured) at the same chip cost.

2. The candidates

Each: the paragraph, the known-failed test as a harness run, the chip row, the verdict. The harness names: v6inv-census.sh is /srv/builds/v6-invention-census.sh on build-1 (a copy sits in this lane's scratch and lands under tools/attack/v6-invention/ with the 09:00 report), the counter-asic-4 crate's igneum-pow accept --class <form> over seeds igneum-v6inv/<i>; the F8 census is tools/attack/f8-uniform (attack-f8 warps, master); the chip row is the model's arithmetic with the GPU side measured where a pack ran.

2.1 Layer 5: the shadow placed per load, one pass of a long sub-block (KEEP, rank 1)

The class v4 shadow runs 256 instructions 27 times after instruction 63 of every iteration, where the next iteration's 64 base instructions and 16 loads stand between the block and every load, and a chip may run it on a second die with 64 B of lane state crossing once per iteration (140 GB/s at the 5090's read rate, a PCB link; 16.2). Placed inside every read's dependency the same work makes the lane state cross twice per read, 2.2 TB/s, an interposer-class link or one N5 die carrying controller, lanes and PHY, which the mission lane's model prices at a project of about USD 60 M against 30 M and a break-even cap of about 200 M against 100 M. The 16 x 27 form (16 sub-blocks of 16, each iterated 27 times) was built on 7 October and closed the same night: 27 passes of a 16-instruction map right before a load collapses or biases the load's address register before any base instruction can re-randomise it, and the acceptance rule in execution order refused it. The sound form named there and never drawn is one pass of a long segment per load: 16 sub-blocks of 432 instructions, each run once, the same 6,912 instructions per iteration. The experimental class caps the per-load block at 4,096 instructions, so today's census takes the two forms inside the cap that bracket it: mx8+shl4096x1 (16 x 256, one pass, 59 percent of class v4's work per hash) and mx8+shl2304x3 (16 x 144, three passes, class v4's exact count), and the constant-work ladder between them (16 x 256 x 1, 32 x 128 x 2 ... 16 x 16 x 16) that reads how the acceptance rate depends on the sub-block length against the pass count. The 432 x 1 form itself needs the cap raised, a one-line change in a research-only parser, which this lane does not make this week; its number is bracketed by the 256 x 1 and 144 x 3 rows and the ladder's flatness between them.

Known-failed test, as a harness run: ERA=none SEEDS=256 /srv/builds/v6-invention-census.sh on build-1 under lease pool 48. Fail: a form whose per-candidate rejection is 0.98 or above (P(exhaust) at the 256-attempt cap above 0.5 percent, an epoch without a program every few months). The 16 x 27 form reads 0.989 and fails (section 3.1). Pass: a form under 0.95 (P(exhaust) under 2 x 10^-6). The 256 x 1 form reads 0.927 and passes; so does every form with a sub-block of 36 instructions or longer (0.927 to 0.949). The second known-failed case is the instrument's own: ERA=era on the same seeds reads 0 of 256 accepted on every per-load form, with the window's fixed top bits named as biased index bits (section 3.3); the fixed instrument (bits at or above 28 - k_off_s excluded from the value-level test) must accept per-load forms under drawn eras at the no-era rate within the binomial band, and still refuse the per-load record of 7 October (candidate 0 of the first export, 1,482 duplicate lanes).

Chip row:

Column Value Label
A fixed-function chip's k unchanged (the work is the same ALU mix); its capex: the controller cannot be a 28 nm part with the core elsewhere, so the project moves from USD 5 M (no core) or 30 M (a core on its own die) to about 60 M (one N5-class die or a 2.5D package); break-even cap about USD 200 M in years 1 to 2 (s = 0.30) against about 100 M modelled (16.2, the mission lane's N3 single-die row; a GDDR7 PHY on N5 is unpriced)
A GPU-like chip's per-joule edge unchanged: 2.1x at k = 1, 3.5x at k = 0.3 at the knee; its capex per MH/s 3.0 to 4.3 USD against 2.8 modelled
RTX 5090 the sound forms (PC 1, 12:14 to 12:20 UK): shl2304x3 (class v4's count) 135.85 MH/s at 483.6 W unlocked against the whole block's 137.62 at 464.6 (+19 W, -1.3 percent of rate; the premium over mx8 1.09 against 0.95 microjoules), at the 1,300 lock 125.88 at 303.4 against 126.99 at 296.8 (+6.6 W; 0.72 against 0.66 microjoules); shl4096x1 (59 percent of the count) 135.99 at 428.6 W unlocked and 126.08 at 271.8 locked (0.68 and 0.47 microjoules over mx8). Per shadow instruction: 19.7 and 20.8 pJ unlocked against the whole block's 17.2, so the placement costs this card 15 to 20 percent more per instruction. The 16 x 27 v2 export (dead as a class) had read 13 to 14 W UNDER the whole block; the sign reverses on the sound form, which is the number the row carries measured (the hash lane's job; the mx8 control 137.65 at 312.2 W and 127.39 at 212.6)
RTX 5070 Ti no row (no card in this lane); the 4070's class v4 premium at its tune point, 30 W for no rate, less the block effect, is the proxy approximate
Apple M5 Max the 16 x 27 v2 export 26.88 MH/s against 27.01 (-0.5 percent, Metal packbench, 7 October); the 256 x 1 form's inline text is 4,096 shadow lines per iteration where the 1,024-line block cost the M5 Max 17 percent (6 October, measured), so its footprint is the open Apple number: OWED (a Mac measurement under the measure lock, which this lane does not run; the hash lane's or the shipper's Metal row) measured for 16 x 27; the sound form's row owed
The verifier 8.28 to 8.55 ms cold with the sibling loaded (256 x 1), 8.44 to 8.82 (144 x 3), against class v4's 8.33 to 8.63 on the same core in the same minutes; the acceptance's dynamic test 35 ms per candidate (1,111 ms for 32), 13.7 candidates per seed on average: about 0.5 s of one core per epoch measured (section 3.2)
The acceptance 0.927 per candidate (256 x 1), P(256 consecutive rejections) 0.927^256 about 4 x 10^-9 per seed; class v4 sub-version 3 reads 0.681 and 2 x 10^-43; the gap is the missing dataflow rule (the per-load class is not the class v4 shape, so (a'), (c') and (c'') do not run on it; the value-level bias test catches the same population: 1,563 of 2,329 bias rejections name index bit 0 at a one-count near 4,096 or 12,288 of 16,384, the product's low-bit law) measured (section 3.1)

Verdict: KEEP as layer 5, the first thing v6 adds beyond the four, because it is the only mechanism found that moves the project cost (about 2x on the mission lane's model) against a measured cost to the 5090 of 4 percent more watts at the same work unlocked (19 W) and 2 percent at the knee (6.6 W), 1.3 percent of rate: the rule of this file (keep what raises the chip's capex more than every GPU tier's cost) holds by 2x against 2 to 4 percent, but the honest line is that the card pays, not saves, and the Apple footprint is still unread; its acceptance is now a measured 0.927 with a named fix (the sub-version 3 dataflow fixpoint run over the real execution order, base and sub-blocks interleaved, as the generator's draw rule) that the research file's 20.2 already asked for. What it does not do: move the energy identity by one joule. The founder's "useless as soon as it dropped" is layer 1's and 3's sentence; layer 5's sentence is "the chip that can be built costs twice as much to start".

The hash runs on 8 registers per lane, a prototype value to be fixed at gate 1 (spec 1.4). The lane state a chip must carry is those 8 words plus the nonce and counter, 64 B, which is why the chip's 1,172 lanes are 73 KB of SRAM and why, under layer 5, the per-read crossing is 128 B at 17.5 G reads per second, 2.2 TB/s: an interposer carries that (a one-stack HBM package moves about 0.8 to 1.2 TB/s of memory traffic and a die-to-die link of a few TB/s is a 2.5D product; approximate, from memory), so the chip has an escape at USD 200 of package instead of one die. Draw the register count per era from {8, 16, 32} (the acceptance rule's part (b) over every register; the program length scaled so that every register is written, or registers above 8 initialised and read by the shadow alone) and the crossing is 128 to 256 B per lane per read: 4.5 to 9 TB/s, past the interposer class, so the single die is forced and the project's USD 60 M row has no cheaper branch. A GPU pays nothing in rate while latency-bound: a CUDA lane holds up to 255 registers, the 5090's 7,262 lanes in flight at 256 B are 1.9 MB against about 43 MB of register file (170 SMs x 256 KB; approximate), and occupancy at 32 live registers plus the kernel's temporaries fits the 64K-register SM at full residency (approximate, unmeasured). The verifier's register-major arrays grow 4x (4 KB per warp) and the interpreter's cost per op does not move.

Known-failed test, as a harness run: the generator with REGISTERS as a class field (a research-only change behind a pack name, mx8+r32), then v6inv-census.sh over the per-load forms at 8, 16 and 32 registers. Fail: the 32-register form's per-candidate rejection above the 8-register form's by more than the binomial band (more registers, more cold registers, more (b) rejections unless the program length scales). Pass: rejection at or under the 8-register form's; the F8 census at 2^24 on 64 seeds within 1.2x of the window model (the top 0.1 percent of items); the cold verify on core 40 with core 88 loaded under 10 ms. Then the card: one pack per register count on the 5090, rate within 1 percent of the 8-register pack at both states (the occupancy claim measured, not argued).

Chip row:

Column Value Label
A fixed-function chip with layer 5: the lane state per read 128 to 256 B, 4.5 to 9 TB/s of die-to-die traffic at the 5090's read rate; the interposer branch (USD 200 of package) closed, the single N5-class die forced; the project about USD 60 M either way, but with no cheaper escape; alone (without layer 5): nothing, the state crosses once per iteration modelled, the link figures approximate
A GPU-like chip's per-joule edge unchanged; its lane SRAM 73 KB to 300 KB (nothing) modelled
RTX 5090 / 5070 Ti / M5 Max 0 rate while latency-bound (approximate: the occupancy arithmetic above; a measurement is the pack); watts: the same F (the same ops) approximate until the pack runs
The verifier 4x the register arrays per warp (4 KB); cost per op unchanged (0.1 ns per lane-instruction, the shadow's law) modelled
The acceptance part (b) over 16 or 32 registers needs the base program to write every register: at 64 instructions over 32 registers about 13 percent of registers are never written (approximate, e^(-64 x 0.75 / 32)), so either the base length scales with the register count (the verifier's 10 ms gate holds to about 330,000 ops) or the extra registers belong to the shadow alone and part (b) reads the base's 8 modelled; the census decides

Verdict: KEEP as layer 6, conditional on layer 5 (alone it moves nothing). Its value is one sentence in the chip's project plan: no interposer saves the second die.

2.3 Layer 7: warp-uniform data-dependent block selection (KEEP, rank 3, small)

A program whose control flow depends on the data it reads is the brief's first candidate. Per-lane branches are dead on arrival: a divergent branch costs a GPU warp both paths and a chip with per-lane sequencers nothing (the history's "placed nowhere" table; RandomX's one predictable branch targets speculative CPUs, which Igneum does not have). The form that survives is warp-uniform: at the end of each iteration a value reduced across the 32 lanes by shuffles (xor-fold of r0, say, which costs 5 shuffles) selects which of B drawn sub-blocks runs next, the same block for every lane of the warp, so the GPU takes one uniform indirect branch per iteration (as sel already takes one per iteration for the immediates) and the FPGA overlay or the fixed pipeline must hold all B blocks and pay B times the shadow's area for one block's throughput. For the GPU-like chip, a sequencer that already runs the hour's program, it is one more jump. What it buys: the per-epoch bitstream (the FPGA lane, history addition 5) holds B blocks instead of one, so a mid-size part's compile (42 to 160 minutes, PRflow, claimed in spec 1.13.1) carries B times the logic; at B = 4 a part that fitted one block does not fit, and at B = 8 the overlay must time-multiplex. Nothing in the energy identity moves. The acceptance must run every reachable path (B blocks per iteration, each judged by (a') in its own order) and the uniformity census must read the selection's bias (a selection that favours one block is a block that runs more).

Known-failed test, as a harness run: a research-only class mx8+sh256x27+sel<B> (the generator draws B blocks, the interpreter selects per iteration from the warp-folded r0), then igneum-pow accept over 256 seeds with every path judged, and attack-f8 warps at 2^24 on 64 seeds. Fail: the block-selection histogram over the 2^24 nonces outside 6 sigma of uniform (a plant: select from lane 0's r0 bit 0 alone, which the fold is meant to prevent), or any path's (a') verdict differing from the whole-program verdict. Pass: within the band, 60 of 64 seeds under 1.2x on the hot-set test, as class v4 reads.

Chip row:

Column Value Label
A fixed-function chip or an FPGA overlay B x the shadow's logic for one block's throughput, or time-multiplexing at 1/B the rate; a bitstream compiled per epoch carries B blocks approximate (LUT area scales with the straight-line block; no FPGA row exists in the repo)
A GPU-like chip nothing: one jump per iteration on a sequencer modelled
RTX 5090 / 5070 Ti / M5 Max about 0: one uniform branch per iteration, 5 shuffles per iteration for the fold (5 x 8 = 40 shuffles per hash at 29 to 56 pJ: 0.002 microjoules, 0.3 percent of F); the compile-ahead carries B x 256 instructions of text (B = 4: the 1,024-line footprint that cost the M5 Max 17 percent on 6 October) modelled on measured per-op costs; the Apple footprint is the cap on B
The verifier B x the acceptance's dynamic test per candidate (every path); the hash's cost unchanged modelled

Verdict: KEEP, rank 3, with B capped by the Apple footprint (B = 2 or 4 at the 256-instruction block, or B = 4 at 64-instruction blocks, which the 6 October measurement says run 2.5 to 3.5 percent faster anyway). It is the only candidate that moves the FPGA lane, which the history ranks as the first adversary of a per-hour program (Lyra2REv2, X16R) and which no measured row in the repo has priced (the HBM FPGA row is 0.30x to 0.39x of a 5090 per watt, chip-model-v3 5.3, the soft-overlay case unmeasured).

2.4 The reserve ordered by hardware orthogonality (KEEP as an ordering rule inside layer 3)

Layer 3 unlocks reserve families by height and rotates after exhaustion. The order is Open in spec 1.13.2 except R1. The history's addition 6 said: families that force a full 32-bit datapath per lane first, the int8 tile last. Today's measured rows (research file 15.1a; the research lane's layer 3 table) sharpen it: a shuffle crossbar is the one block where the GPU's cost per op is highest (29 to 56 pJ) and a chip's is near it (about 20 pJ, approximate), so shfla (lane plus delta, a second crossbar form) is the family a 12-op chip lacks most and gains least on; byte permute and popcount next (small adders a chip adds for 0.1 pJ, but a datapath without them loses 4 points of the mix); the int8 tile last (k 0.03 to 0.3: a chip's MAC array is cheaper than the GPU's tensor core, so the tile is kept for datapath diversity and never for joules). The chip row is layer 3's: about USD 4 of N5 pre-wires all eight. Cost to the cards at 4 points: under 1 percent of ALU time on every vendor (measured step costs: shfla 1.91x on Apple, 1.53x NVIDIA, 0.75 to 0.84 AMD). Known-failed test: the layer 3 gate's own (a kernel built without the live family refused at packcheck; the fast-time harness crossing one family epoch with a stale miner, 0 accepted blocks after it). Verdict: not a new layer; an ordering rule, stated so the synthesis fixes it at genesis.

2.5 Data-dependent program graphs, per-lane (REJECT)

The brief's form: the program's control flow drawn from the data it reads, per lane. A GPU warp executes a divergent branch as both paths with lanes masked, so a branch taken by half the lanes doubles the ALU work of that span; a chip with a sequencer per lane pays the taken path only. Number: a shadow of 55,296 instructions per hash with one two-way branch per 64 instructions at 50 percent divergence costs the card up to 2x the shadow's premium (165 W instead of 83 at the knee on the 5090, modelled on the measured 6.4 pJ per op) for a chip cost of 1x; k on the branched work falls to 0.15 to 0.4. Costs GPUs more than chips. The warp-uniform form (2.3) is what survives.

2.6 Latency-bound reads tied to the shard proof (REJECT)

The hash's reads sampled from the state the miner is proving: class v5 already keys every item to a leaf of the execution state at the epoch's cut and refreshes per epoch (spec 1.8.6; proof of following). Tying the reads to the segment being proved means a refresh per block (every second) from the touched leaves. The research file's row 7 priced the refresh as a cost: a 1 GiB rebuild is 157 G ops, 13.4 ms on a 5090 and 32 ms on a 4090 (measured), 0.16 to 0.5 J on a chip core (1 to 3 pJ per op, approximate); per block that is 0.3 to 1 W against 78 W of chip hashing (0.4 to 1.3 percent) and 1.3 to 3.2 percent of a GPU's hash time (the rebuild stalls the hash on the card; the chip's rebuild runs on its core beside the memory). A delta refresh (only the touched leaves, a few KB) costs both sides nothing. Either way the GPU pays more or equal. What the tie would buy is liveness (a chip must follow the chain per block, not per epoch), which class v5's per-epoch refresh already gives at the WAN line of 2a.2. Number: GPU 1.3 to 3.2 percent of rate against a chip's 0.4 to 1.3 percent of energy. Rejected.

2.7 Randomised memory topology per era (REJECT)

The dataset's address map and stride family drawn per era: class v3 draws the stride multiplier M, the rotation R and the interleave pos per era already (spec 1.13.1, Counter ASIC 2.0 layers 4 and 8, decided IN at a six-era hash-rate spread of 1.3 percent on the 5090, 3.2 on the 9070 XT, 0.8 on the M5 Max, measured). The plan said then what still holds: a chip whose address decoder can permute its address lines pays nothing. Drawing a richer family (a per-era permutation polynomial over bank and row bits, a drawn item size, a drawn line interleave across devices) costs the chip's decoder a few hundred gates and costs the honest card whatever the mapping does to its own DRAM's bank parallelism: a mapping that concentrates consecutive dependent reads into one bank group hurts the side with fewer lanes in flight, which is the chip (1,172 against 7,262), but the chip adds lanes at 64 B each (lane state is free, chip-model-v3 5.5), so the asymmetry closes at no cost. The one topology lever that would have moved the chip, the hot region above the window model, was read by adv-cache-2 as the diffuse era-stride excess (a product's low bits placed at address bit R; 8 of 27 drawn-era programs over 1.2x), which is a FAULT the next class's value-level test removes, not a lever to keep. Number: 0 to the chip, 0 to 3.2 percent to the cards. Rejected; the existing draws stand.

2.8 VRAM-size ratchet (REJECT as a lever; layer 2's floor stands)

A floor that rises with chain state by rule is layer 2. The ratchet form (the floor tracking the modal miner's VRAM minus the prover footprint, or rising on a calendar faster than the schedule) was read against the card-lifetime table (docs/analysis/card-lifetime-2026-10-05.md, option (b) steps): the 4 GB tier ends at the 4 GiB step, 8 GB at 8 GiB, 12 and 16 GB at the 16 GiB step; a chip holds 24 GB (one HBM3 stack, about USD 200, modelled) or 32 GB (the 5090's own 16 devices, USD 320), so every step retires a card tier before it touches the chip, and at 32 GiB and beyond the chip adds devices and its activate-bound rate RISES with the bank count (chip-model-v3 5.7, row "dataset size": "not a lever against this chip"). Number: at the 16 GiB step the 8, 12 and 16 GB tiers are out (3 of 6 card tiers) and the chip's energy per hash moves 0. Rejected as a lever; layer 2's rule (the schedule as the floor, the state above it, a ceiling at the next cache doubling) is kept exactly as the synthesis writes it, with its honest line that it retires cards before chips.

2.9 Proof-carrying hashes sampled by the pool (REJECT)

A fraction of hashes carries a verifiable execution witness. Three witness forms were read. (i) The hash's own 128 item values: the stored-dataset chip has every item; the recompute chip derives them; cost 0 to both, 512 B per share on the wire. (ii) A Merkle witness of the state leaves under the window's root: the chip's node has it (one node serves a farm, class-v5 2a.2); cost 0 to both. (iii) A witness that the item was DERIVED (a transcript of the 8 dependent cache reads and the mixer's 72 applications): a stored-dataset chip cannot produce it without the cache and the mixer core, so this form forces the f = 0 chip's silicon (the 256 MiB SRAM mirror, USD 46 of die and an N5 project) onto the f = 1 chip for the sampled fraction g; but the honest GPU must produce the same transcript, and deriving an item on the card is 8 dependent 64-byte cache reads at the mixer's 9,360 ops (the inline kernel measured 4.8x slower than the honest kernel on the M5 Max, spec 1.8.5), so at g = 1/128 (one item per hash) the card pays about 4 percent of its rate and at g = 1/16 about 30 percent; the chip derives on an SRAM-resident cache at 6.3 nJ per item (modelled) for 7 percent of its energy at g = 1/16. Number: GPU 4 to 30 percent of rate against the chip's 1 to 7 percent of energy. Rejected on form (iii); forms (i) and (ii) force nothing.

2.10 Prover-gated eligibility (REJECT)

The block's eligibility tied to the miner's proving (a key must have proved its share of segments in the last window to claim a block), so a chip farm must carry provers. The bound is the gas bound the research file's section 7 found: the chain needs 2 shards per block at the v1 budget, about 1.6 kW of 5090 proving network-wide at 1 block per second (measured prover rows), independent of the hash rate. Against the hash: at 1 GH/s the hash draws 2.4 kW (2.4 microjoules per hash, measured), so proving is 67 percent of it; at 100 GH/s 0.7 percent; at 10 TH/s 0.007 percent. A chip farm at share s must prove share s of 1.6 kW: six 5090s per farm at any scale, which is the node it already runs. Redundant proving (each segment proved by m provers) raises the forcing m times and is the useful-work gaming the history records (Aleo, Boundless). Number: at mainnet scale the forcing is under 0.01 percent of the chip's energy; the honest card already proves. Rejected; the 80/20 split stands.

2.11 Time-locked parameter commitments (REJECT beyond the era VDF)

An era's parameters committed under a VDF so a chip cannot be built ahead: the era VDF of 7 October (docs/analysis/era-vdf-2026-10-07.md) already makes the era draw's input unknowable for 517 s on the fastest prover measured (chiavdf's GMP path, 208,800 squarings per second, against the production T of 108 million) and the era lead is 2 hours (pow_era_lead), so a chip taped out against era n knows era n + 1's draw 2 hours before it runs, against a 5-month (Bitmain) to 13-month (a startup) design cycle (the history's lesson 5, Vorick). Lengthening the delay or the commitment changes nothing a chip can use: the GPU-like chip holds the whole drawn band as firmware (the synthesis's section 7), and the fixed-function chip is dead at the first draw outside its wired value whether it learns the draw 2 hours or 2 days ahead. The one party for whom 2 hours matters is the FPGA fleet (a bitstream compiles in 42 to 160 minutes on a mid-size part, claimed), and layer 7 (2.3) and the epoch length (spec 1.13.1, the 600 s floor) are the levers for it, not the lock. Cost of a longer lock: one honest node core for the VDF's hour per era (today) rising linearly with T; the 2019-class verify gate already missed by 2.2x (26 ms against 10, measured, proxy). Number: 0 to the chip at any delay above 2 hours; the honest node's core-hours rise with T. Rejected.

2.12 A fraction of reads derived from the cache in the hash (REJECT)

The brief's spirit of "tie the hash to what the chip must hold": a fraction g of the 128 reads per hash derived on the fly from the 256 MiB cache (8 dependent cache reads and 72 mixer applications) instead of read from the dataset, so the stored-dataset chip must carry the recompute chip's cache and core for that fraction. This is 2.9 form (iii) without the witness and the same arithmetic: the card's derived read is 8 dependent DRAM reads (the cache does not fit L2 at 256 MiB, and the cache doubling keeps it so), so at g = 1/16 the card's dependent-read count per hash rises from 128 to 184 and its rate falls about 30 percent (modelled on the latency-bound rule; the inline kernel's 4.8x at g = 1 is the measured anchor); the chip with the cache on die derives at 6.3 nJ per item (4.0 nJ of SRAM reads, 2.3 of mixer; modelled) for 0.466 to 0.50 microjoules per hash (+7 percent) and buys the USD 46 mirror and the N5 project it already needs for the shadow core. Number: GPU -30 percent of rate at g = 1/16, chip +7 percent of energy and +USD 46 of die. Rejected.

2.13 Row-straddling and double-activation reads (REJECT)

The chip and the card share the DRAM's physics (the research file's section 5: the same tRC, the same activate window, the same 32-byte atom). A read that opens two rows (an item straddling a row boundary, or two independent 32-byte sectors per read) costs the chip's memory +0.9 nJ per read (a second 909 pJ activation; modelled) and the card +1 sector of traffic, which at 128 reads per hash is the w64 regime (64 B per read): the 5090 fell to 71.9 MH/s, bandwidth-bound, 47 percent of its rate (measured, read-width). Number: chip +45 percent of E_mem (0.466 to 0.58), card -47 percent of rate and about +9 percent of energy per hash on the memory side alone. The edge moves from 3.6x to about 3.1x at the knee (modelled) at the price of half the card's rate. Rejected.

2.14 Per-era lane-state and scratch draws (REJECT)

Per-lane live state across the hash (a scratch with read-modify-write) was measured out in Counter ASIC 2.0 (layer 3: the recompute chip's gain at every share 2.4x, the cards -12 to -48 percent) and bounded in docs/analysis/scratch-soundness.md (the live state sits in a chip's SRAM at under 5 percent of its mirror). A drawn scratch size per era draws from a dead family. Number: the cards -12 to -48 percent of rate (measured), the chip +picojoules per access. Rejected. (The register-file width of 2.2 is the live form of this idea: state that costs the GPU nothing because its register file is already there, and costs the chip a link, not an SRAM.)

2.15 A refresh per block (REJECT; the research file's row 7)

Dead by arithmetic: a 1 GiB rebuild is 0.16 to 0.5 J on a chip core and 13 to 32 ms of a card's hash time; per block that is 0.4 to 1.3 percent of the chip's energy and 1.3 to 3.2 percent of the card's rate. The refresh cadence is a liveness tool (class v5's proof of following), not an energy lever.

3. The measured rows (igneum-build-1, 8 October 2026, 11:0x to 11:2x UK)

The crate: igneum-pow of branch counter-asic-4 at 5984ffab, built on build-1 through tools/build-remote.sh --no-fetch --box 1 from the detached worktree igneum-wt-v6-inv-ca4 (RESULT rc=0, 11 s, sccache); the binary /srv/builds/igneum-wt-v6-inv-ca4/igneum-pow/target/release/igneum-pow. Every run under /srv/builds/_bin/lease (pool 48 at class measure for the censuses, cores 40,88 for the benches), owner class-v6-invention; the box at load 15 to 24 on 96 threads from other lanes throughout; logs under /srv/builds/v6-invention/ (census-none.tsv, census-era.tsv, logs/<era>-<form>-<seed>.log, census-run-*.log), copied into docs/analysis/class-v6/logs/ with the 09:00 report.

3.1 The acceptance census (v6inv-census.sh, 256 seeds igneum-v6inv/0..255, every candidate's verdict through igneum-pow accept --class <form>, the class's own 32-attempt cap)

No era (ERA=none; 2,816 rows in 80 s on 48 cores):

Form (sub-blocks x length x passes) Shadow instructions per iteration Seeds accepted of 256 Seeds exhausting 32 attempts Candidates Rejection per candidate Mean accepted attempt First failing part, the top four
mx8+sh256x27 (class v4's shape, the whole block after instruction 63) 6,912 256 0 268 0.045 0.05 (a) 6, (b) 5, (c) 1 (the v2/v3 rule only: the crate's accept does not take the class v4 parts on this spelling, so the row is the control's shape, not sub-version 3's 0.681)
mx8+shl256x27 (16 x 16 x 27, the 7 October form) 6,912 76 180 7,043 0.989 15.9 per-load bias 6,081; (c) 545; (b) 212; (a) 129
mx8+shl256x16 (16 x 16 x 16) 4,096 69 187 7,117 0.990 15.4 bias 6,294; (c) 406; (b) 215; (a) 133
mx8+shl576x12 (16 x 36 x 12) 6,912 215 41 3,815 0.944 10.6 bias 3,367; (b) 98; (a) 72; (c) 63
mx8+shl512x8 (16 x 32 x 8) 4,096 209 47 4,120 0.949 11.5 bias 3,686; (b) 106; (a) 82; (c) 37
mx8+shl1152x6 (16 x 72 x 6) 6,912 226 30 3,428 0.934 9.9 bias 3,036; (b) 97; (a) 62; (c) 7
mx8+shl1024x4 (16 x 64 x 4) 4,096 231 25 3,487 0.934 10.6 bias 3,097; (b) 81; (a) 73; (c) 5
mx8+shl2304x3 (16 x 144 x 3, class v4's count) 6,912 224 32 3,457 0.935 9.9 bias 3,076; (b) 89; (a) 67; (c) 1
mx8+shl2048x2 (16 x 128 x 2) 4,096 228 28 3,423 0.933 10.1 bias 3,041; (b) 83; (a) 65; (c) 6
mx8+shl4096x1 (16 x 256 x 1, the sound form inside the cap) 4,096 234 22 3,217 0.927 9.7 bias 2,823; (b) 94; (a) 64; (c) 2
mx8+shl4096x2 (16 x 256 x 2) 8,192 233 23 3,273 0.929 9.9 bias 2,875; (b) 95; (a) 67; (c) 3

What the ladder says: the iterated 16-instruction map is the fault (0.989 to 0.990 whatever its pass count), and from a 36-instruction sub-block up the rejection is flat at 0.927 to 0.949 whatever the pass count or the work per hash (4,096 or 6,912 or 8,192 instructions per iteration). The remaining 0.93 is not the placement: the (c) lane-constant and distinct rejections fall to 1 to 7 per form (they were 406 to 545 on the 16-instruction forms), and the bias rejections are the product's low-bit law at the load's source, which under class v4 the sub-version 3 dataflow rule (a') removes at the draw and which the per-load class, not being the class v4 shape, never applies. The value-level reading of the 256 x 1 form's 2,329 bias rejections: index bit 0 in 1,563 (one-counts clustered at 3,584 to 4,608 of 16,384, the 1/4 law, 680 of them; and at 11,776 to 12,288, the 3/4 complement, 212), bits 26 and 27 in 502 (one-counts 7,168 to 7,680: a mild low bias of the top address bits just outside the 6-sigma band of 384, the high bits of small products through mulhi, unattributed), the other 26 bits 264 in all; by site, site 0 takes 788 of 2,329 (its source is written last by the previous iteration's sub-block 15 and by the base instructions before instruction 1, where the draw's redraw covers the sub-block's last writer of the next load's source and not a product rotated into place by a later rotl). The named fix is one rule, the research file's own ask of 20.2: the sub-version 3 freshness fixpoint and the shared-operand rule run over the real execution order (base instruction, the load, its sub-block, the next base instructions), and a load whose source is not fresh at that point refused at the draw, which class v4 pays at 0.568 of candidates and which should bring the per-load forms from 0.93 toward 0.68.

The acceptance's own cost: 35 ms per per-load candidate on one box core (1,111 ms for 32 candidates; the control's v2/v3 rule 2.8 ms), so an epoch's draw at 13.7 candidates is about 0.5 s of one core; the (c'') ratio at 2^20 would add class v4's 2.8 s per chosen candidate.

3.2 The verifier (core 40 of the EPYC 9454P at nice 19, core 88 its SMT sibling held busy by a 100,000-warp class v4 bench for the whole run, killed by its pid at the end; the ladder's method; igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <form> --warps 50)

Form Cold, sibling loaded, warp base 0 / 4,096 / 1,000,000 (ms) Avg of 50, loaded (ms) Cold, core alone, earlier in the same minutes (ms) Label
mx8+sh256x27 (class v4's shape) 8.63 / 8.33 / 8.33 8.29 5.90 / 6.34 / 5.24 measured; the ladder's run 2 read 8.77 / 5.14 at load 25 on 6 October
mx8+shl4096x1 8.55 / 8.28 / 8.28 8.24 5.44 / 5.21 / 5.17 measured
mx8+shl2304x3 8.82 / 8.54 / 8.44 8.43 5.31 / 5.13 / 5.09 measured
mx8+shl4096x2 not run loaded 5.97 / 5.33 / 5.29 measured, alone only

The per-load forms verify in class v4's time within the run's noise (the same instruction count, the same items derived, 4,096 per warp on every row). A first pass of the loaded column was discarded: its sibling run (300 warps) ended before the measured warps began, the ladder script's own known-failed case, and read the quiet-core figures; the second pass held the sibling for the whole run.

3.3 The instrument fault under drawn eras, and the correction it forces

The same 2,816 rows with --era igneum-era-test/<seed mod 16> (ERA=era): every per-load form 0 of 256 seeds accepted, 8,192 candidates per form, rejection 1.000; the class v4 shape 256 of 256. Of the 256 x 1 form's 7,862 bias rejections, 5,536 name index bit 26 or 27 "set in 0 of 16,384" or "set in 16,384 of 16,384" at a load site, which is the era's working-set window (spec 1.13.1: k_off = below(3) per site puts the site on the whole dataset, a half or a quarter by fixing the top k bits of the index to the drawn offset o); the prototype's BiasedIndexBit test loops bits 0 to 27 and does not exclude bits at or above 28 - k_off_s, so under any era it refuses every program with a half- or quarter-window site, which is nearly every program. The research file's 20.2a-close ("the 16 x 27 per-load class accepts 22 of 1,621 candidates over 64 seeds, 1.4 percent; 42 of 64 seeds exhaust the chain's 32 attempts") was read across drawn eras and therefore measured the instrument on most of its rows; the no-era census above is the construction's own figure (the 16 x 27 form 0.989 per candidate, 76 of 256 seeds accepted, which still fails the test of 2.1 and keeps that form dead). Owed to the research file (the Counter ASIC coordinator): the instrument fix (skip the window bits per site) and the 20.2a-close figures re-read with it; neither is made this week by this lane, which changes no code in the crate.

3.5 The uniformity read (tools/attack/v6-invention/uniform, 11:5x to 12:0x UK; 64 seeds igneum-v6inv/0..63, the first accepted candidate of each, 2^20 nonces, every load's index through the library's trace_load_indices on the closed-form dataset at 2^28 words, no era; per site the distinct-index ratio d_s / E_s at N = 2^23 evaluations, the cross-hash item histogram's top 0.1 percent share against a uniform SplitMix64 control of the same read count; logs docs/analysis/class-v6/logs/uniform-*.tsv)

Form Seeds read (exhausted at the 32 cap) Min site ratio: min / median / max Sites under 0.995 (the (c''') floor) Top 0.1 percent share against the control: min / median / max Seeds over 1.2x Heaviest item, reads (the control's max 29) Reading
mx8+sh256x27, drawn through the v2/v3 rule only (the crate's accept on this spelling runs no (a'), (c') or (c''): the class v4 shape BEFORE sub-version 3) 64 (0) 0.093 / 0.540 / 1.00008 60 0.999 / 7.42 / 36.3 41 396,564 (seed 39, site 13), 358,161, 350,025 AP-F8-1's class seen whole: the value-level hot sets the dataflow rule and the floor were added for; the harness's known-failed case fired on real programs, which is what makes the two rows below a reading
mx8+shl4096x1 (16 x 256 x 1, the sound form) 60 (4) 0.99707 / 1.00009 / 1.00013 0 0.9984 / 1.0003 / 1.0339 0 (1 over 1.02x) 251 (seed 23, site 13; 8.7x the control's max, 0.0002 percent of the reads) clean on 59 of 60 at the F8 line; one mild hot item on seed 23 that the 0.995 floor admits (0.99707) and the top-0.1-percent share reads at 1.034x: the residual class of the in-house pass (a few-item concentration under the floor's resolution, about 1.0004x to a chip there, under 1.002x here)
mx8+shl2304x3 (16 x 144 x 3, class v4's instruction count) 61 (3) 0.99962 / 1.00009 / 1.00013 0 0.9984 / 1.0002 / 1.0019 0 (0 over 1.02x) 31 clean on every seed read: the uniform control's own figures

What the three rows say together: the per-load forms, with their draw's redraw rule (a sub-block writer of the next load's source never a product or an or) and the value-level bias test, read as clean at 2^20 as the uniform control on 119 of 121 seeds, where the raw class v4 shape through the same rule reads hot on 41 of 64; the chain's class v4 sub-version 3 (the dataflow fixpoint and the (c'') floor) read 60 of 64 under 1.2x on F8's census at 2^24. So the two per-load tests are at least as strong as sub-version 3's on this read, at the price of the 0.93 rejection of 3.1; and the dataflow rule in execution order, when it lands, is expected to move the rejection down and the uniformity not at all. Owed: the same read under drawn eras on the fixed instrument (the research lane's commit), and at 2^24 on the four seeds nearest the line (23, 29, 49 and the first exhausted).

The box record, stated because it cost the network something: the 24-thread run of this harness on build-1 took 30.8 to 31.7 GB resident (16 hash sets of 2^20 indices and two 2^24-bucket histograms per thread, about 1.2 GB per thread) and the kernel's OOM killer ended its second and third forms at 11:58 and 11:59 UK; the same pressure killed the Devnet 3 seed and node1-dn3 on build-1 in that window (the build-server lane's read, 12:0x UK). The lease holds cores, not memory; the rerun was 8 threads (about 10 GB, released after 684 s with 0 kills); every further run of this lane names its resident memory in the lease label and goes to build-4 (88 GB free at 12:0x UK), never the seed box. The build-server lane carries the memory rule for the lease tool and the hands' oom_score_adj.

3.4 The packs

igneum-pow export --seed igneum-genesis --day 2026-10-03 --class <form> --out <dir> on build-1: mx8_shl4096x1 (attempt 7, id 75ca9547da21b200, OVERALL PASS, kernel.cu 245,267 bytes, 4,308 lines) and mx8_shl2304x3 (attempt 8, id bbfdfc1dcdda0b46, OVERALL PASS, kernel.cu 142,010 bytes, 2,516 lines), under /srv/builds/v6-invention/packs/, tarred as v6inv-perload-packs.tgz (sha256 ca1986b7fc5fab20a643fc37151a55e01f91edbfacc6f1a22a7384ae87cc11bc). Handed to the hash lane at 11:1x UK and run in its job run-ca4-pc1-v6-packs-5090-20261008 (PC 1, the RTX 5090 alone, 12:14 to 12:20 UK, 250 batches of 2^24 per row, nvidia-smi at 1 Hz, the helper's cleared lock sequence; self-test PASS and the pinned fingerprint on every row: shl4096x1 180edd33ae32a4ae, shl2304x3 a209664a9b3b7bf1; raw rows node tools/jobs.mjs run-ca4-pc1-v6-packs-5090-20261008 --all). Nothing of the packs' output goes to a served page or the spec.

Pack Shadow instructions per hash Unlocked (2,842 to 2,865 MHz): MH/s / W / MH per W Over mx8, unlocked At the 1,300 MHz lock: MH/s / W / MH per W Over mx8, locked pJ per shadow instruction, unlocked / locked
mx8-genesis (class v3, the control) 0 137.65 / 312.2 / 0.441 127.39 / 212.6 / 0.599
mx8_sh256x27 (class v4's shape, the whole block) 55,296 137.62 / 464.6 / 0.296 152.4 W, 0.95 microjoules 126.99 / 296.8 / 0.428 84.2 W, 0.66 17.2 / 12.0
mx8_shl2304x3 (the sound per-load form at class v4's count) 55,296 135.85 / 483.6 / 0.281 171.4 W, 1.09 125.88 / 303.4 / 0.415 90.8 W, 0.72 19.7 / 13.0
mx8_shl4096x1 (one pass of 256 per load) 32,768 135.99 / 428.6 / 0.317 116.4 W, 0.68 126.08 / 271.8 / 0.464 59.2 W, 0.47 20.8 / 14.3

What the rows say: on this card the per-load placement pays 15 to 20 percent more per shadow instruction than the whole block at either state (19.7 to 20.8 pJ against 17.2 unlocked; 13.0 to 14.3 against 12.0 locked), with the rate 1.1 to 1.3 percent under the whole block's. The 16 x 27 v2 export's 13 to 14 W saving (research file 20.3) does not carry to the sound form; that export was a 16-instruction block iterated 27 times, a loop the compiler keeps in registers, where the sound form is a 144- or 256-instruction straight segment per load. The chip side does not move with any of this (k is the chip's cost per op over the card's at the same work), so the column that changes is the honest card's: layer 5 costs a 5090 19 W unlocked and 6.6 W at the knee over class v4 for the project-cost doubling it buys.

4. The ranked list, every column

Rank Layer Door (section 1) Chip: fixed-function k / capex Chip: GPU-like per-joule edge Cost: RTX 5090 Cost: RTX 5070 Ti Cost: Apple M5 Max Verifier Harness state
1 Layer 5: the shadow per load, one pass of a long sub-block capex and project (door 2) k unchanged; project about USD 60 M against 30 M, cap about 200 M against 100 M (modelled); capex 3.0 to 4.3 USD per MH/s unchanged (2.1x at k = 1) +19 W unlocked, +6.6 W at the knee against class v4's block at the same work, -1.3 percent of rate (the sound form, measured) no row; the 4070's 30 W premium as the proxy (approximate) -0.5 percent (16 x 27 form, measured); the 4,096-line footprint OWED 8.28 to 8.82 ms loaded (measured) acceptance 0.927 measured; the dataflow rule in execution order is the fix (gate plan)
2 Layer 6: the register-file width drawn per era capex (door 2), with layer 5 closes the interposer branch: 4.5 to 9 TB/s of die-to-die traffic (modelled, approximate link figures) unchanged about 0 rate (occupancy arithmetic, approximate; the pack measures it) the same the same 4x the register arrays, cost per op unchanged (modelled) no pack form; a generator constant (gate plan)
3 Layer 7: warp-uniform block selection life of a bitstream (door 3, the FPGA lane) B x the shadow's area for a fixed pipeline or an overlay (approximate) unchanged about 0 (one uniform branch and 40 shuffles per hash: 0.3 percent of F) the same the compile footprint caps B (the 1,024-line block cost 17 percent, measured) B x the dynamic test per candidate no pack form (gate plan)
4 The reserve's order life of fixed silicon (door 3) USD 4 pre-wired unchanged under 1 percent at 4 points (measured steps) the same the same 0 layer 3's gate

What the list does not contain, and why it is honest to say so: no candidate moves the per-joule identity. The chip that stores the dataset keeps 3.6x at zero premium and 2.1x at k = 1 under every layer here, as under the four. What the list adds is the project cost (rank 1 doubles it on the mission lane's model, rank 2 removes its cheaper branch), the FPGA lane's cost (rank 3), and the order of the reserve (rank 4). The founder's line "render an ASIC useless as soon as it dropped" is true of the fixed-function chip under layers 1 and 3 and stays false of the GPU-like chip under everything; the honest public sentence is the research file's: the price per joule of the honest card's operating point and the shadow's premium are what hold the general chip, and the layers decide which chip can be built and what it costs to start.

5. Consequences per tier (the standing rule of 5 October 2026)

Tier What this file means today What is being done
A home miner, one 8, 12 or 16 GB card, any vendor, any OS nothing changes: no layer here ships this week, no class moves, the devnet pays nothing; if rank 1 lands in v6 the card runs the same shadow work in a different place at the same or lower watts (the 16 x 27 form's measured 13 to 14 W under the whole block on a 5090; a 4070-class card's premium at its tune point is 30 W today, measured) the sound forms' 5090 rows today; a small-card row (the 4070) by job when the hash lane's queue allows
One 24 or 32 GB card (5090 class) the measured rows of section 3; the sound per-load placement costs this card 2 to 4 percent more watts than class v4's block at the same work and 1.3 percent of rate; a 5090 owner under layer 5 at the knee pays about 7 W more than under class v4 the rows are in; the knee tune (rank 1 of the research file) is the lever that pays for it
RTX 5070 Ti no card in the project; every row is the 4070's or the 5090's scaled, approximate; the hash lane's default (the stock pair plus the 5090's lock slope) stands stated as approximate wherever it appears
Apple (M-series) the one open number: the inline footprint of a 4,096-line per-load block (the 1,024-line block cost 17 percent on 6 October); if it costs rate, rank 1 takes a block-shape cap for the Apple tier (144 x 3 at 2,516 lines, or a 64-instruction sub-block form) and the synthesis says so a Metal packbench row under the measure lock by the lane that runs the Mac (not this one)
A rig watts per card as the 5090 row; a rig's bill under rank 1 is at or under class v4's the same rows
A pool user nothing: no share, payout or template changes in any candidate kept; the rejected 2.9 (pool-sampled witnesses) is the only one that would have touched the pool protocol nothing
A node operator (the verifier) the per-load forms verify in class v4's time (8.3 to 8.8 ms loaded, measured); layer 7 at B blocks multiplies the acceptance's dynamic test per candidate, not the hash the 2019-class core measurement (O-1.14) decides rung and block caps as before
A chip maker under rank 1 the controller and the shadow core share one advanced die or an interposer (project about USD 60 M, modelled); under rank 2 the interposer no longer suffices; under rank 3 a per-epoch bitstream carries B blocks; the per-joule edge is unchanged the gate plan of section 6
The public claim nothing moves; the chip texts rest on the research file's close (2.1x at k = 1 for 82 to 90 W on a 5090 at the knee, measured four times) the Counter lane's texts

6. The gate plan (hours, never weeks; nothing this week)

Gate What runs Pass line Known-failed case
G5-draw (layer 5) the generator's dataflow fixpoint and shared-operand rule over the real per-load order; the 4,096 cap raised so 16 x 432 x 1 draws; v6inv-census.sh at 256 seeds, no era and drawn eras (with the instrument's window bits excluded) rejection per candidate at or under 0.75 on every sound form, 0 exhaustions, drawn-era rate equal to the no-era rate within the binomial band; the F8 census at 2^24 on 64 seeds: 60 of 64 under 1.2x, the four tail seeds' sites read against their own windows the 7 October per-load record (candidate 0 of 854050a4293f0615) refused; the 16 x 27 form at 0.989 refused by the line
G5-card (layer 5) one pack per sound form on the 5090 (today), the 4070 and the 9070 XT by job, Metal packbench on the M5 Max rate within 1 percent of class v4's shape at both states on NVIDIA and AMD; watts at or under class v4's; the Apple footprint within 5 percent or the block-shape cap set a pack whose fingerprint differs from the Rust verifier's on any vendor
G6 (layer 6) the register count as a class field; the census at 8, 16, 32; the card packs 2.2's pass line 2.2's fail line
G7 (layer 7) the selection class; every path judged; the selection histogram at 2^24 2.3's pass line the lane-0 plant
G-order (the reserve) layer 3's gate with the order fixed layer 3's line layer 3's case

Hours: the dataflow rule in execution order 2 to 3 (the research file's own estimate) plus the re-export and census 1; the register class 3 to 4; the selection class 4 to 6; the card jobs are queue time. Nothing is coded this week; the first code is the founder's call after the synthesis.

7. Unverified and owed

  • The uniformity read under drawn eras (the fixed instrument's commit), and at 2^24 on the seeds nearest the line.
  • The sound forms' card rows on the 4070 and the 9070 XT (the 5090 rows are in, section 3.4); the M5 Max footprint by the lane that runs the Mac.
  • The 16 x 432 x 1 form itself: bracketed by the 256 x 1 and 144 x 3 rows (0.927 and 0.935) and the flat ladder between 36 and 256; not drawn (the cap).
  • The instrument fix (the window bits) and the re-read of 20.2a-close: owed to the research file's owner, not made here.
  • The bits 26 and 27 mild bias (502 rejections at one-counts 7,168 to 7,680 of 16,384): unattributed; a trace of the site's source writers is the next read.
  • The register-width and block-selection classes: modelled only; no pack exists.
  • Every chip figure is the model's (chip-model-v3 section 5 and the research file's sections 2, 15.1a, 16 and 20.4); the die-to-die link figures of 2.2 are approximate, from memory, uncited; no chip has been measured.
  • The web search budget of this session was exhausted before this lane's reading; every external figure here is one already cited in the repo's files, with its URL and date there (the research file's section 14, the history's section 6, chip-model-v3 5.1).

8. Sources

Internal: docs/spec/01-lottery-hash.md (1.4.3 to 1.4.7, 1.8.5, 1.8.6 on branch class-v5, 1.13); docs/analysis/chip-model-v3.md (sections 1 to 3, 5.1 to 5.11, 6); docs/analysis/counter-asic-4-research.md on branch counter-asic-4 (sections 0, 2, 4, 7, 9, 15.1a, 15.1b, 16, 17, 20.2 to 20.4); docs/plans/cryptanalysis/in-house-pass.md (sections 12 to 14); docs/plans/counter-asic-3-status.md section 7c; docs/analysis/asic-resistance-history.md (sections 1.2, 2.4 to 2.6, 3, 4.3); docs/analysis/era-vdf-2026-10-07.md; docs/design/class-v5-stored-state.md on branch class-v5 (2a, 3); docs/design/latency-ladder.md on branch ladder (2 to 5); docs/analysis/card-lifetime-2026-10-05.md; docs/analysis/latency-shadow-2026-10-06.md; docs/analysis/scratch-soundness.md; docs/design/class-v6-rotating-family.md on branch counter-asic-4 (the synthesis's outline at 5984ffab); the harness runs of section 3 (logs on build-1 under /srv/builds/v6-invention/).

External, as cited in those files (read there on the dates they state): O'Connor et al., Fine-Grained DRAM, MICRO 2017; Li, Reddy, Jacob, MEMSYS 2018; Folded Banks, ISCA 2025; Horowitz, ISSCC 2014; Dally, Hot Chips 2023; the mlsysbook energy table citing Horowitz and Dally; the RandomX design document; the ProgPoW audits (Least Authority and Rao, 2019); Condrey, PoSME, arXiv 2604.15751 (April 2026); the Ethash, RandomX, Equihash and Cuckatoo chip rows of the history; the GDDR7 price rows (TrendForce, September 2026).