docs/analysis/horizon/new-pow.md sections 0 to 9: scheme A (mining is proving) never, on bytes, the verifier and sampleability; scheme B (the tensor-shaped integer shadow) prototyped as proto-newpow/mma-shadow and measured, never as class content on the energy reading, with the R8 two-output correction; scheme C (proof of stored state, sd1: the daily dataset derived from the execution state) prototyped as proto-newpow/state-dataset, measured on the GPU and the box's CPU, and put forward as the class v5 candidate with its spec items and the Devnet 2 gate. The lane's standing rule: a shadow lever only works through joules the honest card is forced to spend, so shadow work goes where the GPU is least efficient per op. Chip rows in sim/horizon/new-pow/chip_rows.py by the chip-model-v3 method. Rented box addresses replaced by placeholders in the READMEs and the run script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
78 KiB
Horizon lane 8: a new proof of work (three candidate schemes, reviews, prototypes, verdicts)
6 October 2026, evening UK, lane new-proof-of-work, worktree /Users/joshm/Projects/igneum-wt-horizon (branch horizon from master, at 3f4f719). Output of this lane: this file and proto-newpow/<scheme>/. Nothing here touches the shipped hash, igneum-pow, the node, the manifest or the live devnet; every prototype is a benchmark beside the worker, never inside it.
the project lead's mandate, verbatim: "if we create a new way of hashing or a new way of proof of work to revolutionise the space then that's absolutely fine, I want you to deploy everything to create something that has not ever been done before."
0. Progress (kept current for the coordinator)
| Time (UTC) | State |
|---|---|
| 19:05 | Lane started. Read: preamble, CLAUDE.md, the two personas, spec 01 (whole), 04, 07, chip-model-v3 (whole), asic-resistance-history (sections 0 to 3 and 4.3, 5), latency-shadow-2026-10-06 (whole), counter-asic-3-status sections 1 to 5, counter-asic-3-node section 6 (the P2 signalling rule), int8-matrix-family sections 1 to 3, scratch-soundness verdict, proving-methods (whole), fud-ledger M1, M7, M16, M22, M28, P2, F13, proto-cuda host.cu and the mx8-genesis pack (kernel.cu, memhard.h, program.h, vectors.h), proto-cuda/emu, family-probe.cu, the fleet's prover-tiers-real-cards.md, bench-log line 2582 (rental cost) |
| 19:25 | Fleet agent asked for two boxes; answered at 19:29: two quiet RTX 4090s (RunPod, nvcc 12.8 at /usr/local/cuda/bin, directory /root/horizon-newpow, until 22:30Z). No quiet Ampere card exists tonight; a loaded 3090 is offered. Main's note: cost rows use bench-log 2582 (USD 0.0117 per MH/s-hour) |
| 19:35 | File skeleton written. Two prototype sub-agents launched (budget two at once): mma-shadow on box 1 (), state-dataset on box 2 () plus CPU rows on igneum-build-1. Designs being written in this file meanwhile |
| 19:41 | Section 3 complete: the three designs, the one-table comparison, the migration path. Scheme A's verdict is already visible in its own numbers (A1 dead on 2.9 MB of openings per block, A2 dead on sampleability and a 32 to 40 ms proof verify; A0 is scheme C with the trace as state). Prototypes running: mma-shadow (box 1) and state-dataset (box 2 and igneum-build-1). mm8 two-output correction sent to the prototype |
| 19:50 | Section 4 (reviews, the pick: B and C), section 7 (ranked next steps) and section 8 (open questions) written. Scheme C prototype complete and measured on box 2 and igneum-build-1: hash rate and watts unchanged (63.08 vs 63.09 MH/s, 207 W), build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact 1,024 of 1,024 items and 128 of 128 lanes; rows in 5.2. Scheme B ladder running on box 1 (R = 0, 8, 32, 128 measured, 512 in progress) |
| 20:20 | Scheme B ladder complete on box 1 (R = 0 to 512, rate flat at 63.08 MH/s, 201 to 216 W, every fingerprint PTX = reference, 1,024 of 1,024 lanes at every R, verifier delta 0.05 to 4.39 ms). Sections 5.1, 5.3, 6 and 9 written. Box 2 released 19:53Z; box 1 released on the mma agent's report. File complete |
1. What was read and the facts this lane stands on
Every figure below is from the named file; "approximate" marks a figure from memory.
| Fact | Value | Source |
|---|---|---|
| The shipped hash | 64 instructions x 8 iterations, 16 loads per program (128 dependent 4-byte reads per hash), 8 registers, 32-lane unit with xor shuffles, class v3 = mixer x8 item derivation over a 256 MiB ChaCha12 cache, 1 GiB dataset in the packs (2 GiB designed), era draws, VDF seeds | docs/spec/01-lottery-hash.md 1.4 to 1.13 |
| Verifier today | 2.06 ms per unit on one M5 Max core (class v3, 4,096 item derivations), about 5.2 ms on a 2019-class core by the 2.5x rule; the 10 ms gate | docs/plans/counter-asic-3-status.md section 3, latency-shadow-2026-10-06.md section 4 |
| RTX 5090 at the hash | 136.1 MH/s (readwidth), 132.2 (shadow control), 290 W in the app, 350 W in the bench; 2.34 to 2.65 microjoules per hash; 17.5 G dependent reads per second, 82 percent of the GDDR7 activate ceiling; 45.2 T int op/s; marginal ALU energy 10 to 13 pJ per counted op | chip-model-v3.md 5.1, latency-shadow-2026-10-06.md 5 |
| The chip that matters | the f = 1 stored-dataset memory-controller chip: 5.1x per joule on GDDR7, 7.5x to 9.2x on HBM3 in the model; 2.1x to 4.8x by the Ethash precedent; the recompute chip (f = 0) 0.31x per chip, 1.86x per joule | chip-model-v3.md 5.4 to 5.6 |
| The one lever against it | program work in the latency shadow: at N = 100,000 ops per hash the chip's edge over the 5090 falls from 5.6x to 2.1x at k = 1 (chip core energy per op equal to the GPU's 11 pJ), to 3.2x at k = 0.5, 4.1x at k = 0.3; class v4 candidate mx8+sh256x27 |
latency-shadow-2026-10-06.md 6 and 10 |
| Step costs per family on the 5090 (ratio to the add-xor-rotate chain, 7,941 G lane-steps/s) | rotr 1.32, shflx 1.49, shfla 1.53, dot4 1.16, mm8 (mma.m8n8k16.u8, bit-exact) 2.43 |
counter-asic-3-status.md section 3, item 6 |
| mm8 on other vendors | AMD RDNA 4: WMMA iu8 builtin reaches gfx12, 1.68 to 1.83 per step, fragment layout UNVERIFIED (exactness not attempted); Apple: no integer simdgroup matrix in MSL, Metal 4 matmul2d uchar x uchar into int exists but not from the Swift toolchain used here, per-lane dot4 emulation 1.6x unsigned, 4.7x signed |
counter-asic-3-status.md item 6 AMD column, int8-matrix-family.md 1 and 4 |
| Proving today | SP1 6.8.1 Hypercube; the v1 shard (4.7 M cycles) proves in 4.8 to 18 s on 12 to 32 GB cards with the patched server, 7.4 to 8.0 GB alone; compressed proof 1.27 MB, verified in 32 to 40 ms; the aggregator 2.2 to 9.7 s per block | prover-tiers-real-cards.md, proving-methods.md 1 and 2.1 |
| Proof payment | 80/20 lottery/proving split of the subsidy; shards by weighted sortition (8 assignees, 10 s window), aggregator share 1,000 bps, unproven deadline 600 DAA s; the native-execution veto: a record whose statement differs from the node's own execution pays nothing | docs/spec/07-execution.md 7.2, 7.7, 7.8 |
| Class activation | P2: a class flips when 95 percent of blue blocks over a one-day window carry the object byte (header version high byte), with a fixed-height floor; one-sweep binary rollout; Devnet 2 gate first | docs/plans/counter-asic-3-node.md section 6, CLAUDE.md 6 Oct rules |
| Rented hash | USD 0.0117 per MH/s-hour (1,748 MH/s for USD 20.44 per hour on RunPod community pods, 18:45Z); the live devnet 1.16 GH/s | docs/bench-log.md line 2582 |
2. Method
Designs first (section 3), each reviewed in two personas (section 4), two picked, two prototypes measured on real cards (section 5), verdicts (section 6). Prototype shape: the mx8-genesis pack's own kernel text (proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cu, memhard.h) modified by the smallest change each scheme needs, compiled with nvcc 12.8 on the two RunPod 4090 boxes, timed by CUDA events over 2^24-nonce batches with nvidia-smi at 1 Hz, and checked bit for bit against a C CPU reference that interprets a 32-lane unit register-major with lazy item derivation (the verifier's shape) on 1,024 random lanes. CPU rows on igneum-build-1 (EPYC 9454P, Zen 4, one core at up to 3.8 GHz, which is NOT a 2019 core; the 2.5x rule of the project stands in for that core, labelled). Chip rows are the chip-model-v3 method (section 5 of that file) applied to each scheme, approximate where that file is approximate.
3. The three candidate schemes
Shared notation: K_d the day key (spec 1.8.1), S_e the epoch seed words (spec 1.3), the unit = 32 aligned nonces (spec 1.9), item(t) the 16-word class v3 derivation of spec 1.8.5, target64 as spec 1.10. Difficulty in every scheme is the unchanged 64-bit target comparison on the unchanged DAA (spec 2.3): none of the three changes what a block's work unit is worth, only what the unit of work consists of, so the difficulty controller sees the same statistics. Each scheme is a program class in the sense of spec 1.4.5 (a generator number and a class byte), so activation runs through P2 (section 3.5).
3.1 Scheme A: mining is proving ("proof of committed trace")
The claim to test. The lottery's work is a bounded piece of the chain's own proving, so the 80/20 split collapses into one payment and the hash rate is the proving capacity.
The puzzle, in its most favourable form. Every full node already runs the segment natively (spec 7, the native-execution veto). Add one step to that: every node also runs the zkVM executor (SP1's RISC-V executor, CPU, no proving) on the segment's shard inputs and keeps the shard TRACE: the cells of every table the shard touched, about 280 M cells for the adopted v1 shard, 1.1 GB (proving-methods.md 1.4). The lottery of epoch e then uses as its dataset the trace of the last segment whose last chain block has DAA score at most 3,600 e - 1,200 (the same 20-minute lead as the epoch seed, spec 4.3), serialised row-major and padded by zero to the dataset size, and the item derivation becomes item_A(t) = class v3 derivation with s[i] ^= row(t)[i] for the 16 words of trace row t (64 bytes of trace per item). Everything else is the shipped hash: 128 dependent 4-byte reads, the 32-lane unit, the fold, target64. Header commitment: nothing new; the trace is a function of the segment, the segment of the chain, the chain of the header's past, exactly as the epoch seed is (spec 4.3 item 4); the class byte 5 of P2 marks the object. Difficulty: unchanged. Verifier: the node derives up to 4,096 items per unit from the 256 MiB cache (as today) plus 4,096 reads of the trace rows it holds in RAM (1.1 GB): the measured cost of exactly this read pattern is scheme C's row in section 5 (the two schemes share the verifier shape). Under 10 ms on a 2019 core if scheme C's row is.
Two stronger forms, and why each dies on a number.
| Form | What the miner must hold or do | Verifier | Why it dies |
|---|---|---|---|
A1, "proof of committed codeword": the dataset is the Reed-Solomon codeword of the trace (BaseFold's stacked encoding at blowup 4, 4.5 GB, proving-methods.md 1.1 and 1.4), whose Merkle root the shard record already publishes; the miner must have done the COMMIT stage of the proof (encode and hash) to mine |
the LDE and the Merkle tree: the first stage of every STARK or BaseFold prover, the part a GPU spends a large share of its proving time on (approximate: 30 to 50 percent of the stage time, unmeasured here) | cannot derive a codeword word locally: one evaluation of a stacked column at one point is O(2^21) field operations (the stacking height, proving-methods.md 1.1), so 4,096 words per unit is about 8 G operations, about 1 s on a core, 100x over the gate. The block must therefore carry Merkle openings: 4,096 per unit (every lane's loads feed every other lane through shfl) x 22 levels x 32 bytes = 2.9 MB per block; with a 1-lane unit (no shfl) 128 x 22 x 32 = 90 KB per block, 7.8 GB per day of PoW witness at one block a second against about 200 B per header today (Kaspa header, approximate). Headers-first validation, the pruning proof and the light client (spec 10) all break on the bytes |
|
| A2, "mine a proving step": each nonce selects a random piece of the pending proof work (a FRI fold of one column, a Poseidon2 Merkle layer, a zerocheck round) and the hash is of that piece's output | that piece | the verifier either recomputes the piece (then the verifier did the useful work, and the piece bought nothing: the definition of useless) or verifies a proof of it (an SP1 compressed proof verifies in 32 to 40 ms, proving-methods.md 2.1, four times the whole gate, and a per-piece proof does not exist). And the pieces run out: a segment's proof is about 5 s of one 5090 (prover-tiers-real-cards.md: 4.8 to 18 s per shard on 12 to 32 GB cards, 2.2 to 9.7 s of aggregation) against 8 G hashes per 8-s segment at 1 GH/s; the useful fraction of the lottery's work is bounded by (proving work per segment) / (network hashes per segment), about 8 percent at 1 GH/s and 0.08 percent at 100 GH/s (arithmetic on the cited figures), because proving work is set by gas and lottery work by the security budget. They are the same quantity only by coincidence at one network size |
The sampleability objection, stated and answered. Ball, Rosen, Sabin and Vasudevan, "Proofs of Useful Work" (ePrint 2017/203) build PoUW for problems with a random self-reduction (orthogonal vectors, 3SUM, all-pairs shortest paths): a useful instance is embedded into a random challenge so that the challenge is hard on average, solving it solves the instance, and verification is fast; the price is a polynomial blow-up and a problem class with that structure. Ofelimos (Fitzi, Kiayias, Panagiotakos, Russell, CRYPTO 2022) makes the useful work a doubly-parallel local search whose QUALITY improves with more work, so more hash rate yields better solutions. Primecoin (2013) mined Cunningham chains (useless, but sampleable: the nonce picks the chain's origin). Gridcoin pays BOINC credit through a trusted whitelist, not a puzzle (approximate, from memory for the last two). Proof generation has none of the three properties these constructions need: (1) a segment has one proof, not a distribution of instances, so there is nothing for the nonce to sample; (2) partial progress has no verifiable value under the gate without a proof, and a proof verify costs 32 to 40 ms; (3) the quantity of useful work is fixed by demand (gas), the quantity of lottery work by the security budget (hash rate), and a puzzle whose per-solution work is fixed by demand is not a difficulty-adjustable lottery. The answer this design gives is therefore no: the only form that survives the gate and the bytes is A0 above, in which the miner must HOLD the trace, which is "proof of stored state" with the trace as the state, and the useful work gained per hash is zero (every node computes the trace natively anyway; ledger F13). Mining does not become proving; it becomes proof that the miner executes.
Chip edge per joule (chip-model-v3 method). A0 changes the dataset's contents, not its size or read pattern; the f = 1 stored-dataset chip stores whatever the items are: 5.1x per joule on GDDR7, 7.5x to 9.2x on HBM3, unchanged from chip-model-v3.md 5.4. The f = 0 recompute chip must now hold the 1.1 GB trace somewhere (it cannot derive an item without row t), so it needs DRAM beside its SRAM and becomes an f = 1 chip; the 0.31x row disappears, which is a small gain since that chip was never the threat. A1 and A2 are not priced: they fail before a chip is drawn.
Verifier cost. A0: the class v3 derivation (2.06 ms per unit on an M5 Max core, measured) plus 4,096 random 64-byte reads from a 1.1 GB array; the read row is measured in section 5 (scheme C's CPU rows, the same pattern). A1: 2.9 MB of openings or a 1-lane unit (see table). A2: a proof verify (32 to 40 ms, measured) or nothing useful.
Bit-exactness. A0: the trace rows are bytes; the derivation is integer; nothing vendor-specific is added. The zkVM executor's trace must itself be deterministic across platforms (SP1's executor is Rust, no floating point in the trace path, approximate: not audited here); any disagreement about a trace cell is a consensus split, so A0 adds the SP1 executor to the consensus-critical code, which ledger P7 already names as a risk class for the proofs and which here would extend to the lottery.
Known attacks. Grinding: a block producer influences the trace through the transactions it includes, but the trace used is 20 minutes old and keyed by K_d through the mixer; a producer cannot predict a useful bias in 128 dependent reads of a keyed derivation (the MTP lesson, asic-resistance-history.md 2.3: Dinur and Nadler controlled addresses by controlling contents; here the contents enter only through a keyed, chained derivation whose addresses are the cache-line indices of the mixer state, not the trace). Outsourcing: pools serve the dataset (1 to 2 GiB per miner per day), so "the miner holds the trace" becomes "someone in the pool holds it". Precomputation: the dataset is computable 20 minutes early, as today. Light evaluation: as the f = 0 row above. Sampleability: the paragraph above. Empty segments: a day or an epoch with no transactions has a near-empty trace (the shard statement still applies rewards, spec 7.7 item 8), so the dataset is mostly padding and the hash degrades to today's class v3: graceful, not an attack.
Game theory of the collapsed payment (if A1 or A2 had worked). The 20 percent pool would go: provers are miners and the proof is the by-product of the lottery. Who is paid for what: the block reward pays for the block and the proof piece in it; a prover that mines earns exactly a miner that proves. Pools: the pool does the useful work and sells shares, as today, so proving centralises exactly as mining does (Aleo's proof of succinct work centralised on the same path: the coinbase puzzle was a synthetic circuit proven by whoever had the most GPUs; approximate, from memory, the cryptographer persona's own list). Proving demand at zero: the puzzle has no useful content and must fall back to a synthetic dataset, so the chain carries two puzzle modes and a mode switch, a consensus rule and a grinding surface. External jobs: a customer's proof could only be mined if its segment entered the puzzle, so the security budget would be spent on the customer's work at the marginal cost of inclusion, a subsidy from holders to customers unless priced by a burned fee. These are the reasons CLAUDE.md records the lottery and the proving as separate on purpose; nothing found tonight overturns them.
Migration. A0 is a class byte (P2): class v5 by the 95 percent signal with the floor height, one-sweep binary rollout, Devnet 2 gate first (CLAUDE.md 6 Oct). Every node must run the zkVM executor on every segment before the flip (CPU cost: the v1 shard is 4.7 M cycles; SP1's executor runs at tens of MHz on a CPU, approximate, so under a second per 8-s segment); a node without it cannot validate PoW after the flip, which is what the floor height is for.
Verdict line (section 6 has the reasoning): NEVER as "mining is proving" (A1, A2); A0 is scheme C with the trace as the state and is folded into C.
3.2 Scheme B: a GPU-structure-bound puzzle ("mx8+mm8xR", the tensor-shaped integer shadow)
The claim to test. Fill the latency shadow (the only lever that moves the f = 1 chip, chip-model-v3.md 5.7, latency-shadow-2026-10-06.md) with work shaped like a GPU's own tensor datapath rather than its scalar ALUs, so that a chip must carry GPU-class matrix units whose energy per operation NVIDIA's own silicon already sits near the floor of, and the chip's residual edge k (its energy per op divided by the GPU's) cannot fall far under 1. The brief's other candidates are rejected on definition: shared-memory bank timing and register-file width are timings and capacities, not values, and a consensus rule can only check values; a per-warp scratchpad was measured and found not sound as a chip layer (docs/analysis/scratch-soundness.md: the live state is bounded by the read-modify-write count, 64 to 320 bytes per lane, which a chip keeps in SRAM at under 5 percent of its mirror).
The puzzle. Class v3 (mixer x8, 16 loads, the 64 base instructions, the era draws) plus a block of R mm8 steps executed at the end of every iteration, after instruction 63 and before the next iteration samples sel; 8 R steps per hash, no load in the block, the base program untouched (the same seam the class v4 shadow uses, latency-shadow-2026-10-06.md section 2, so the acceptance rule of 1.4.6 keeps its verdict draw for draw). Step k has four draws from the program stream after the base and shadow draws: a = below(8), b = below(7) + (b >= a), c = below(8), c2 = below(7) + (c2 >= c). Its semantics over the unit:
A: 8 x 16 uint8 lane l holds A[l >> 2][4 (l & 3) .. 4 (l & 3) + 3] = the 4 bytes of r[a] (byte 0 = lowest k)
B: 16 x 8 uint8 lane l holds B[4 (l & 3) .. +3][l >> 2] = the 4 bytes of r[b]
C = A x B C[i][j] = sum over k of A[i][k] B[k][j], exact in int32 (at most 1,040,400)
r[c] = r[c] + C[l >> 2][2 (l & 3)] (mod 2^32)
r[c2] = r[c2] + C[l >> 2][2 (l & 3) + 1] (mod 2^32)
This is the fragment layout of PTX mma.sync.aligned.m8n8k16.row.col.s32.u8.u8.s32 with the two .s32 outputs d0, d1 of each lane (PTX ISA, "Matrix Fragments for mma.m8n8k16", the integer layout; docs/analysis/int8-matrix-family.md 2.2 quotes it and proto-cuda/family-probe.cu carries the CPU reference that was bit-exact on the RTX 5090 on 5 October 2026, counter-asic-3-status.md item 6). The spec defines the matrices and the lane ownership, not the instruction: a vendor permutes its own fragment layout into this one (a shuffle) and the result is defined whatever the hardware. Why TWO destinations where the reserve entry R8 of 1.13.2 has one: with one element per lane only 32 of the 64 products of C are consumed, so a chip does 512 multiply-adds per step where the GPU does 1,024 and the GPU hands the chip a free 2x on the block; with both outputs consumed the whole tile is load-bearing. This is a correction to the reserve text as well (section 7). Unit of work: still the unit. Header, difficulty: unchanged; the class byte 5.
Parameters and the verifier's law. Per unit the verifier adds 8 R x 1,024 unsigned byte multiply-adds (the register-major interpreter already holds the 32 lanes' registers, so A and B are in hand). Plain scalar C: about 2 ops per multiply-add, 16,400 ops per step per unit; at the 18 G op/s the x8 verifier shows on the M5 Max core (chip-model-v3.md 5.7, measured) that is 0.9 microseconds per step per unit: R = 128 adds about 0.9 ms, R = 512 about 3.7 ms. With a byte-dot instruction (AVX-VNNI vpdpbusd, 64 multiply-adds per instruction; NEON udot, 16 per instruction) the same work is 4x to 16x fewer instructions (approximate). The headroom is the class v3 margin: 7.9 ms steady on the M5 Max core, 4.6 ms on a 2019-class core by the 2.5x rule (latency-shadow-2026-10-06.md section 4), so scalar R is bounded near 500 on the 2019 core and the measured rows of section 5 set it. On the GPU side the 5090's dependent-chain probe runs mm8 at 2.43x the add-xor-rotate step, 7,941 / 2.43 G lane-steps per second = 102 G mm8 per second per card, 104 T multiply-adds per second (the chain probe, not the tensor peak); at 4.25 M units per second (136 MH/s) the card could hide about 24,000 mm8 per hash before the probe rate binds, R about 3,000, far above what the verifier allows. So the verifier binds first, at about R = 500 scalar on a 2019 core (approximate until section 5) and perhaps 4x higher with VNNI: the honest card never leaves the latency bound on the candidate R.
Chip edge per joule (chip-model-v3 method). The f = 1 chip (memory, controller, static: 0.466 microjoules per hash on GDDR7, 0.321 on one HBM3 stack, chip-model-v3.md 5.4) plus a tensor array: energy per hash = memory + 8 R x 1,024 x e_mac x k_mma, where e_mac is the honest card's marginal energy per multiply-add at the block (measured in section 5 as watts delta over MACs per second) and k_mma the chip's ratio to it. The difference from the ALU shadow (k down to about 0.3 for a fixed-datapath array at N5, latency-shadow-2026-10-06.md 6) is where the honest card's engine sits: NVIDIA's tensor cores are int8 multiply-add arrays at N4/N5-class density already, so a chip's array is the same circuit (approximate: int8 MAC datapath about 0.05 to 0.1 pJ at N5, the movement of fragments through the register file the larger term on both sides; from memory) and k_mma is near 1 with a floor near 0.5 for a chip that keeps the fragments in a local register file instead of the GPU's banked one. Rows at k_mma = 1, 0.5 and the measured e_mac are filled in section 5.3 from the 4090 measurement; the shape of the result is already clear: the block lowers the chip's edge by the same mechanism as the ALU shadow and the chip's best case is better bounded, at the price that the honest card's own watts rise by the block's energy (the tensor path is efficient, so the rise per unit of chip-forcing work should be smaller than the ALU shadow's 11 pJ per op; the measurement says).
Bit-exactness across vendors.
| Vendor | Path | State |
|---|---|---|
| NVIDIA sm_75 and later (Turing, Ampere, Ada, Blackwell) | one mma.sync.m8n8k16.u8 per step, identity permutation, wrap on the .s32 accumulate (the wrap edge vector of the R8 entry was bit-exact on the 5090, int8-matrix-family.md 3) |
measured bit-exact on the 5090 (5 October) and on the 4090 tonight (section 5) |
| NVIDIA sm_61 to sm_72 (Pascal GTX 10 series, Volta) | no mma with .u8; emulation: 8 __shfl_sync gathers plus 8 dp4a per lane per step (dp4a.u32.u32, sm_61+) |
unmeasured; cost about 16 steps per mm8 against 2.43 native, approximate; within the 8x emulation bound of 1.13.2 on a per-op basis only if the shuffles are cheap; a GTX 1080 owner is the first NVIDIA tier to pay |
| AMD RDNA 3 and 4 (gfx11, gfx12) | V_WMMA_I32_16X16X16_IU8 with the 8 x 16 and 16 x 8 tiles zero-padded to 16 x 16 and a fixed lane permutation (ds_bpermute) into the spec layout; the RDNA 4 builtin takes 2 ints per lane for A and B (int8-matrix-family.md 1) |
the fragment layout is UNVERIFIED (status item 6: not in any source at hand; the CPU reference was not attempted rather than guessed). This is the gate for AMD: a PC 1 job on the 9070 XT with the spec reference. Until it passes, AMD runs the emulation path (v_dot4_u32_u8 is native on gfx11 and gfx12, 1.06x per op measured) at about 12 to 16 steps per mm8, approximate |
| AMD RDNA 2 and older, CDNA | no WMMA on RDNA 2; v_dot4_i32_i8 exists (dot1-insts, signed only, approximate); CDNA 3 has V_MFMA_I32_16X16X32_I8 with its own layout |
emulation; unmeasured |
| Apple (M-series, Metal) | no integer simdgroup_matrix in MSL; Metal 4 mpp::tensor_ops::matmul2d has uchar x uchar -> int (table 7.3, OS 26.4) through a tensor API not reachable from the Swift toolchain this project uses, and its wrap semantics are unverified (int8-matrix-family.md 1). The emulation: 4 shuffles plus 4 unsigned dot4 emulations at 1.6x per op (measured 5 October), about 10 ALU steps per mm8 |
an Apple miner pays about 4x NVIDIA's per-step cost on the block (approximate); the M5 Max is latency-bound to about 130,000 counted ops per hash (measured), so R = 128 (1,024 mm8 per hash, about 10,000 steps emulated) fits inside its shadow and R = 512 (about 41,000) still does by the ops count, with the hash-rate cost owed to a Metal measurement. Apple's path is the honest card's worst and the chip's argument does not depend on it |
| Intel Arc | XMX through cl_intel_subgroup_matrix_multiply_accumulate (approximate, unverified); dp4a-class dot otherwise |
unmeasured |
So the fleet splits by generation: Turing-and-later NVIDIA and RDNA 3-and-later AMD run the block natively; everything older and Apple emulate. The 5 percent rule (counter-asic-2-public.md) is checked per card with the block live, as 1.13.2 requires for an emulating vendor.
Known attacks. Grinding: none new; the block has no data-dependent control flow and reads no memory. Outsourcing, precomputation: as today (the draws are public per epoch; there is nothing to precompute because the inputs are the per-nonce registers). Light evaluation (the f = 0 chip): untouched; the block never reads the dataset. MTP-class content control: not applicable (no attacker-chosen memory). Sampleability: not applicable. New: (i) the half-tile shortcut, closed by the two-destination form above; (ii) a zero or low-entropy fragment (if r[a] is zero in every lane the step is free): the acceptance rule's register-saturation test (1.4.6 (c)) already rejects programs with stuck registers, and the base program's loads re-randomise every register every iteration; (iii) the trailing-step contraction: the last mm8 of the last iteration writes two registers that feed only the fold; a chip could skip nothing because the fold reads all eight, but a step whose destinations are both never read again before the fold still costs the GPU a full tile; the draw rule should forbid c, c2 outside the fold's register set, which is every register, so there is no such step. (iv) The licensable-IP objection (the history's reason for ranking mm8 last in the reserve, asic-resistance-history.md 4.3 row 6): a chip maker licenses an int8 MMA block at any node. True, and it is the point of the design: the chip must then carry a GPU-class tensor array per 32 lanes in flight at the memory's activate ceiling, 1,172 lanes on GDDR7 (chip-model-v3.md 5.5), 37 tiles in flight, and its edge is k_mma, a ratio of two copies of the same circuit; the design does not claim the chip cannot be built, it claims the chip is a GPU.
Game theory. None of the payment changes; this is a hash change. A pool user sees nothing. A prover that mines pays the block's watts on the same card it proves on; the tensor path is also the proving path's (SP1's Poseidon2 and NTT kernels are integer, not tensor, so there is no contention beyond the power limit, approximate). When proving demand is zero nothing changes.
Migration. Class v5 by P2: the object byte 5, the 95 percent one-day window, the floor height; one-sweep binary rollout because the digest flips; Devnet 2 gate first. The kernel emitter (igneum-pow/src/emit.rs) gains the block in its three dialects, the CUDA one with inline PTX and the IGNEUM_MM8_REF fallback for sm_61 to sm_72, the OpenCL one with the AMD builtin behind a feature test and the emulation otherwise, the Metal one with the emulation; the one-click workers compile the text as they do today (NVRTC accepts inline PTX). Gates before a cut: the six gates of counter-asic-2-rollout.md section 7 plus the AMD layout verification and the Metal emulation's hash-rate cost on the M5 Max.
3.3 Scheme C: proof of stored state ("sd1", the dataset is the chain)
The claim to test. The dataset is the recent chain state, so every hash proves the miner holds the chain, and the lottery's reads double as a verifiable random sample of state for light clients.
The puzzle. Day d's snapshot SS(d) is the execution state at the state root R_d of the certified checkpoint C_day(d), the highest-index checkpoint whose block has DAA score at most 86,400 d - 1,200 (the epoch seed's lead, spec 4.3). Its leaves: the (key, value) pairs of the execution state trie in key order (storage slots, account records, 64-byte chunks of code), leaf t serialised as leaf(t) = Blake2b-512(R_d || t_le32 || key_t || value_t) (the chain's own hash, spec 0.6), 64 bytes each, so every leaf carries full entropy whatever its content and no two days share a leaf; for t beyond the state's leaf count leaf(t) = 0. When the state has more leaves than the dataset has items, the dataset holds the first 2^(D-4) leaves in the order of Blake2b-256(K_d || key): a keyed sample that cannot be chosen without the whole state. The item derivation is class v3's with one line added before the round loop: s[i] ^= leaf(t)[i] for i in 0..15. The hash kernel is byte for byte the shipped one; only the daily build changes. Header commitment: none new. R_d is the state root of a block in the header's own past at a fixed blue score, so "which snapshot was this block mined under" is a function of the header alone once the chain is known, as the program is (spec 1.12). The class byte 5 marks the object. Difficulty: unchanged. Unit of work: unchanged.
The verifier. A node holds the 256 MiB cache (as today) and SS(d) in RAM or mmap (2 GiB at the genesis dataset size, growing on the 1.13.3 schedule), derives up to 4,096 items per unit lazily as today and reads leaf(t) for each: 4,096 random 64-byte reads. The measured cost of those reads and of the derivation is section 5.2. Verifier memory: +2 GiB (+4 GiB at year 4). The daily snapshot build: one pass over the state trie, hashed per leaf (one Blake2b-512 per 64 bytes: about 2^25 hashes for 2 GiB, seconds on a core, section 5.2 measures the stand-in).
What is new, and the prior art. Permacoin (Miller, Juels, Shi, Parno, Katz, IEEE S&P 2014) made the puzzle a proof of retrievability over a large PUBLIC FILE chosen by a dealer, with Merkle openings in each block; the file was external to the chain and the openings were the bytes. Spacemesh (proof of space-time over a plotted file of random data) and Chia (Abusalah, Alwen, Cohen, Khilko, Pietrzak, Reyzin, "Beyond Hellman's time-memory trade-offs with applications to proofs of space", ASIACRYPT 2017; Chia's plots) prove storage of USELESS data. Verthash (Vertcoin, January 2021; asic-resistance-history.md row 8) is the nearest: a 1.2 GB file generated from the chain's own block headers, random reads, no chip after 69 months on a small prize; its data is headers (low entropy per byte, static once written) and it proves nothing about state. Ethash's DAG is from the epoch seed (random). What Igneum C adds: (1) the dataset is the EXECUTION STATE, keyed per day by K_d through the memory-hard derivation, so it is never easier than today's dataset (the leaf is one more 64-byte input to a 9,360-op chain) and it cannot be built from the day key alone: whoever builds it holds the state; (2) every block's 128 loads per lane name 128 keyed items whose leaves are a uniformly random sample of state (the addresses are the mixer-state cache-line indices through 128 dependent reads, unbiasable by the producer at a cost below a block), so a light client that asks any full node for the block's (key, value) leaves with their openings against R_d gets a free daily spot check of state availability; (3) the beacon: the lottery output is already public randomness, biasable by withholding at the cost of a block, as every PoW; C adds nothing there and the design says so. Honest limit: a pool can ship the 2 GiB dataset or the snapshot to its miners once a day (2 GiB per miner per day, 23 MB/s for a thousand miners), so "every miner holds the chain" is really "every mining OPERATION holds the state", which is still a change: today a pool miner needs nothing but the day key.
Chip edge per joule. Unchanged against the f = 1 chip (it stores items whatever they are): 5.1x on GDDR7, 7.5x to 9.2x on HBM3 in the model, 2.1x to 4.8x by the Ethash precedent. The f = 0 recompute chip must hold the leaves (2 GiB) in DRAM to derive anything, so it becomes an f = 1 chip and the 0.31x row disappears. C is therefore not an anti-chip scheme and does not claim to be; it composes with B (the shadow is in the kernel, the state is in the build).
Bit-exactness. The leaf is the output of the chain's own hash over bytes every node agrees on by consensus; the derivation is integer; nothing vendor-specific is added. The one new consensus-critical function is the canonical serialisation of state (key order, chunking of code), which every node must compute identically: a bug there splits the chain on a day boundary, the class of M20 and the DAA 198,000 incident. The Devnet 2 gate exists for exactly this.
Known attacks. State grinding: a producer can write state (pay gas) to influence leaves; the leaf is hashed with R_d, which depends on every leaf, and enters a keyed chained derivation whose read addresses are mixer state, so no bias on 128 dependent reads is reachable at a cost below a block (the MTP lesson, asic-resistance-history.md 2.3, is the reason for the hash and the key, not the plain bytes). Compressible state: an attacker fills state with zeros hoping a chip stores it compressed; the hashed leaves are full-entropy, and the padding region (leaf = 0) is today's dataset, which the chip already stores at 64 B per item. Precomputation: the snapshot is fixed when C_day(d) is certified and K_d is known, 20 minutes before the day (the same lead as the epoch seed), and the build is 13 to 77 ms on the GPUs measured (counter-asic-3-status.md) plus the leaf pass; a reorg across C_day(d) is a merge-depth-scale event, accepted as for the epoch seed (spec 4.3 item 4, O-4.3). Outsourcing: the pool ships the dataset (above). Light evaluation: the f = 0 row above. Long-range: an attacker building an alternative history must build its alternative state snapshots to mine on it, which it does anyway; no change. Finality pause: C_day(d) must be certified; if finality is paused for more than the lead the day's snapshot is not derivable and mining would stop, the coupling spec 4.3 argues against for the epoch seed. Rule, as there: take the selected-chain block at that blue score certified or not, deep enough that a reorg across it is a merge-depth event. Empty state at launch: every leaf is Blake2b(R_d || t || empty), full entropy, the dataset as good as today's: graceful.
Game theory. No payment changes. A miner must run or rent a node (or trust a pool's dataset); the solo-mining floor rises by a full node's state (today's devnet: megabytes; a used chain: gigabytes). A pool user sees a 2 GiB daily download or nothing (the pool serves the dataset). A prover that mines already holds state. A holder gains a daily sample of state availability per block, for free. A rollup customer gains nothing directly.
Migration. Class v5 by P2 (object byte 5, 95 percent over a day, floor height, one-sweep rollout, Devnet 2 first). Every node needs the snapshot builder before the flip (a node without it cannot validate PoW after the flip: the floor height's job). Workers need nothing new: the kernel text is unchanged, the dataset arrives from the node's prepare line as today, built on the GPU from the cache plus a leaf array the node hands over (2 GiB per day over the local socket) or built on the node's CPU and uploaded. The one-click miner's "nothing to install but the driver" line holds; "nothing to download but the day key" does not.
3.4 The three schemes in one table
| Scheme | What it is | Chip edge per joule vs the 5090 (f = 1 chip, chip-model-v3 method) | Verifier ms per unit (model, then measured in section 5) | Vendor bit-exactness | Ships as class v5? |
|---|---|---|---|---|---|
| A, mining is proving | A1 committed codeword, A2 proving steps: dead on bytes and on sampleability; A0 trace-as-dataset survives and is C with the trace as state | A0 unchanged (5.1x GDDR7); A1, A2 not priced | A0: 2.06 + the leaf-read row; A1: 2.9 MB of openings; A2: 32 to 40 ms | A0 adds the zkVM executor to consensus | NEVER as mining = proving; A0 folds into C |
| B, tensor-shaped shadow | class v3 plus 8 R int8 8x8x16 tile steps per hash in the PTX fragment layout, two outputs per lane |
memory + 8 R x 1,024 x e_mac x k_mma; k_mma near 1 with a floor near 0.5 (approximate); rows from the measured e_mac in 5.3 |
2.06 + about 0.9 microseconds per step per unit scalar (R = 128: +0.9 ms; R = 512: +3.7 ms); VNNI 4x to 16x less | native on sm_75+ and RDNA 3+ (AMD layout unverified); emulated on Pascal, RDNA 2, Apple | before measurement: prototype further. After section 5: NEVER as class content (the block costs the honest card 0.056 to 0.70 pJ per multiply-add, so it forces no joules on a chip); kept as reserve R8 evidence |
| C, stored state | the daily dataset derives from the execution state snapshot; kernel unchanged; a state sample per block | unchanged (5.1x GDDR7, 7.5x to 9.2x HBM3); the f = 0 chip disappears | 2.06 + 4,096 leaf reads (section 5.2) | nothing vendor-specific; the serialisation is the consensus risk | before measurement: prototype further. After section 5: SHIP as the class v5 candidate |
3.5 Migration through the class system (common to B and C)
The P2 rule as designed (docs/plans/counter-asic-3-node.md section 6): the header version's high byte carries the producer's object version; epoch e is the new class when the window of one day ending at its seed block has at least 9,500 bps of blue blocks at or above the byte, or when the floor height N is reached, or when epoch e - 1 already was; the rule answers v5 only where it would answer v4. For B and C the object byte is 5 and the floor is set at the publish as DAA + 14,400 rounded up to the epoch boundary. The order: Devnet 2 crossing with tools/fleet/devnet2-gate.sh (zero rejected blocks across the flip, no reorg over depth 3, exec roots agreeing, a segment record paid, every node on the new version), then the live devnet in one binary sweep (the digest flips), then the flip by signal. A node that synced from a pruning proof takes the floor rule for epochs whose window reaches below its pruning point (the same class as the era witness, status item). For C the floor also bounds how long a non-upgraded node can keep validating PoW: none after the flip, so the sweep must be complete before the floor, which is the 10,800-DAA check already in the rule.
4. Reviews: two personas, two short reviews per scheme
Written by this lane in the persona files' voices (.claude/agents/cryptographer.md, .claude/agents/consensus-engineer.md): every design claim with its attack, every claim about another chain with its file, and the smallest change the code allows. Each review names the break if there is one.
4.1 Scheme A, mining is proving
Cryptographer. The break is structural and has a name: a lottery needs a distribution of instances and proving has one instance per segment. A1 moves the work into the commit phase and pays for it in witness bytes (2.9 MB per block, or 90 KB with the unit cut to one lane, which also removes shfl and with it the only thing that makes the 32-lane unit a unit; docs/spec/01-lottery-hash.md 1.9). A2 either recomputes (useless by definition) or verifies a proof (32 to 40 ms measured, proving-methods.md 2.1; the gate is 10 ms). The useful fraction bound (proving work per segment over network hashes per segment) is the argument Ball, Rosen, Sabin and Vasudevan make in the negative direction: without a random self-reduction the embedded instance is a constant, and a constant is amortised to zero by the first miner who computes it. A0 is sound as far as it goes and is scheme C. Second attack on A0 that C does not have: the SP1 executor enters the consensus path for the lottery, so an executor bug that produces a different trace on one platform (an undefined-behaviour corner in a precompile patch, a sha3 or k256 version skew) splits the chain at the PoW, not at the proof; ledger P7 priced that class for proofs where the native veto bounds the damage to one payout; here nothing bounds it. Verdict: never for A1 and A2; A0 only as C with the trace, and then the state is the better choice of data because every node already agrees on it without a second executor.
Consensus engineer. The code says the same thing from the other side. Block validation in the fork validates the header's PoW before the body is fetched and before execution (check_pow on the header path, rusty-kaspa's header_processor; the fork's igneum/exec runs after the block is accepted into the DAG). A puzzle whose verification needs the segment's trace makes header validation wait on the executor of a segment 20 minutes old, which is fine for a synced node and fatal for IBD: a syncing node must execute every segment of history, in order, to validate the headers of history, so headers-first sync and the pruning proof (which validates headers without bodies) are gone; Kaspa's pruning-point sync (consensus/src/pipeline/pruning_processor, approximate location) assumes PoW is a function of the header and a small amount of context. A0 shares this break with C and C answers it in 4.3 by making the snapshot a function of a certified checkpoint's state root plus the state itself, which a pruned node fetches as a snapshot (the exec snapshot path the node already has, CLAUDE.md 6 Oct rule "the p2p snapshot path refuses a snapshot below the node's tip"). For A1 the 90 KB per header kills the header relay and the 600-block record window arithmetic alike. Verdict: never for A1 and A2; A0 is C with a worse data source.
4.2 Scheme B, the tensor-shaped shadow
Cryptographer. No break in the puzzle's soundness: the block is a straight-line integer map with no memory and no data-dependent control flow, its inputs are the per-nonce registers, and the two-output form makes the whole tile load-bearing. Two things to name. (i) The claim that k_mma cannot fall far under 1 rests on the honest card's tensor path being near the floor of int8 multiply-add energy; that is an engineering judgement, approximate, and the external chip review (status item 3, funding.md) is where it is tested, with the ALU shadow's k beside it. The design's advantage over the ALU shadow is bounded, not proven: it narrows the chip's best case from about 0.3 to about 0.5 (approximate) and does not remove the edge. (ii) The fragment layout is a specification of lane ownership; every emulating vendor must reproduce it exactly, and the AMD WMMA layout is unverified (status item 6). Until a 9070 XT run with the spec reference passes, B is bit-exact on one vendor's hardware and in every emulation, which is the state class v2 was in on 3 October (ledger M8) and not a state to cut from. No grinding, outsourcing or precomputation surface is added. A statistical point: the mm8 step's outputs are sums of 16 byte products, so each added value is at most 1,040,400 and its top 12 bits are zero; added into a register it changes the low 20 bits in a structured way. That is fine inside a chain of multiplies and rotates, and the stats run of TESTS.md section 3 should be run on the class before any vector is frozen. Verdict: prototype further; a class v5 candidate after the AMD gate and the stats run.
Consensus engineer. No consensus change beyond the class byte and the emitter: the block is kernel text (igneum-pow/src/emit.rs gains the step in the three dialects), the verifier (verify.rs) gains the step in the register-major loop, the acceptance rule is untouched by construction, and the node's seam is the v5 switch beside the v4 one (counter-asic-3-node.md section 1 lists every file the v4 switch touched; v5 is the same list with one more number). The break to name is operational: the fleet splits by hardware generation. Pascal, Volta, RDNA 2 and every Apple card emulate, and the one-click workers compile three paths where they compile one today; the OpenCL path on AMD needs a feature test at compile time (the WMMA builtin exists on gfx11 and gfx12 and not on gfx10, int8-matrix-family.md 1), which the pack text must carry as a preprocessor branch, and a wrong branch is a wrong hash, which the self-test catches before the worker serves (packfile.h, M28's fix). The 5 percent rule must be checked per emulating card with the block live, and the Apple row is the one most likely to fail it at high R. Verdict: prototype further; set R from the measured rows with the Apple emulation measured on the M5 Max before any cut; the AMD layout is a hard gate.
4.3 Scheme C, proof of stored state
Cryptographer. The construction is never weaker than today's: the leaf is one more input to the same chained derivation, keyed by the day key through the first mixer, and the f = 1 chip's row is unchanged, which the design says. The break to name is the one Permacoin and MTP both met: who controls the data controls the addresses, unless the data enters through a key the controller does not have. Here the controller of state content (anyone paying gas) does not control K_d (a VDF output fixed 20 minutes before the day, spec 4) and the leaf is Blake2b(R_d || t || key || value), so the only lever is choosing content before R_d is known, which affects every leaf through R_d and none in a predictable way. I find no bias at a cost below a block. The second point: the "verifiable sample of state" is real but modest. A block commits to 128 leaves per lane through 128 dependent reads; a light client checking them needs openings against R_d from a full node (depth about 25 at 2^25 leaves, 128 x 25 x 32 B = 102 KB per block if fetched, arithmetic), which is a spot check of availability, not a proof of state correctness; the consensus proof of spec 10 and ledger P4 is still what a light client needs for correctness. The design says this. Third: the canonical serialisation is new consensus-critical code, and the day-boundary flip is the moment it bites (every node rebuilds at once). Verdict: prototype further; a class v5 candidate; the serialisation needs the same test discipline as the DAA switch (a Devnet 2 crossing over a day boundary with a non-trivial state).
Consensus engineer. The break I would have named, the IBD and pruning-proof problem of A0, C answers: the snapshot is a function of R_d, a state root at a certified checkpoint, and the node already carries an exec snapshot path (exec_restart_* fields, the p2p snapshot message, CLAUDE.md 6 Oct). A syncing node validates historical PoW only by holding every day's snapshot, which is 2 GiB per day of history, which is NOT acceptable for IBD. Fix, the smallest I can see: PoW of blocks below the pruning point is not re-validated (Kaspa's pruning proof validates the proof's headers' PoW; rusty-kaspa consensus/src/processes/pruning_proof, approximate), so historical snapshots are needed only for the headers inside the proof, which are a bounded set per level; and for them the proof carries the day's R_d and the node either holds that day's state (recent days) or trusts the certificate (the finality rule's lock already makes those headers irreversible). That is a spec item for docs/spec/10-light-client.md and the pruning section of spec 02, and it is the same shape as the era-seed witness (status item). Second: the verifier's RAM (+2 GiB, +4 GiB at year 4) and the daily build on a node without a GPU (a seed node, a Hetzner box: the leaf pass is CPU work, measured in 5.2) must stay inside the node's budget; the devnet hands on igneum-build-1 have 128 GB, a home node has 16. Third: a day boundary is now a consensus event that depends on a certified checkpoint 20 minutes before it; under a finality pause (6 October, 18:42Z) the day's snapshot falls back to the uncertified selected-chain block, as spec 4.3 argues for the epoch seed; the rule must be written once and tested on the fast-time harness across a pause. Verdict: prototype further; a class v5 candidate; three spec items (pruning-proof witness, the pause rule, the node RAM budget) before a cut.
4.4 The pick
B and C are the two to prototype: both are class objects on the shipped hash, both leave the dataset's memory bound untouched, and they compose (B is kernel text, C is the daily build). A is not prototyped: A1 and A2 fail on bytes and on sampleability before any kernel, and A0 is C with a worse data source. The order of merit at this point, before measurement: C first (no vendor risk, unchanged hash rate by construction, a real new property per block, the pool caveat stated), B second (a real lever against the f = 1 chip with a bounded k, a vendor split and an unverified AMD layout). Section 6 revisits the order on the measured rows: C holds, B retires into the reserve.
5. The prototypes and the measured rows
Both prototypes live under proto-newpow/ with a README carrying the exact commands, the card, the driver and the RESULTS table; this section carries the rows and their consequences. Boxes: two RunPod RTX 4090 24 GB (driver 595.91, nvcc 12.8, -arch=sm_89), quiet, the card to itself; CPU rows on igneum-build-1 (EPYC 9454P, one pinned core, nice -n 19). Power by nvidia-smi at 1 Hz, the mean after the first 10 s of each timed run. Bit-exactness: the PTX path against the reference path on 2^24 lanes (fingerprint), and the GPU against the C CPU reference on 1,024 random lanes.
5.1 mma-shadow (scheme B), box 1
Prototype: proto-newpow/mma-shadow/ (README with every command, kernel_mm8.cu with the PTX path and the IGNEUM_MM8_REF reference path, bench.cu, verify_ref.c, gen_block.py, gen_ref_program.py, run.sh, out/ with every log and csv). Box 1: RTX 4090 24 GB (128 SMs), driver 570.172 (the box reports 570, not the 595 of box 2), nvcc 12.8.93, -arch=sm_89, 1 warp per block, 10 timed batches of 2^24 after a warm-up, power from nvidia-smi at 1 Hz over a 25-s sustained phase (mean after its first 10 s), idle 15.0 to 15.3 W. The two-output tile form of section 3.2 was built (the single-output form never was). Run 19:36 to 19:52 UTC.
| R (steps per iteration) | mm8 per hash (8 R) | MH/s (GPU time) | Ratio to R = 0 | Watts, mean | Block watts over R = 0 | SM MHz | Microjoules per hash | Marginal pJ per multiply-add | Fingerprint (2^24 at base 0) | PTX = reference path | CPU = GPU (1,024 lanes) | Verifier block delta, ms per unit (box core, scalar C) | Registers per thread |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 (control) | 0 | 63.08 | 1.000 | 201.2 | 0 | 2,670 | 3.19 | 7c28cfb06c5c65a9 (the pack's; the 3 pack vectors PASS standalone and in batch) | yes | 1,024 of 1,024 | 0 (the R = 0 run read -0.10, noise) | 29 | |
| 8 | 64 | 63.08 | 1.000 | 204.1 | 2.9 | 2,670 | 3.24 | 0.70 | 06fc2593bfb94b4f | yes | 1,024 of 1,024 | +0.05 | 30 |
| 32 | 256 | 63.08 | 1.000 | 207.7 | 6.5 | 2,670 | 3.29 | 0.39 | 26e83a65f519c865 | yes | 1,024 of 1,024 | +0.17 | 29 |
| 128 | 1,024 | 63.08 | 1.000 | 212.7 | 11.5 | 2,670 | 3.37 | 0.17 | 42223c2113188335 | yes | 1,024 of 1,024 | +1.14 | 32 |
| 512 | 4,096 | 63.08 | 1.000 | 215.9 | 14.7 | 2,670 | 3.42 | 0.056 | 02b7002d747f3711 | yes | 1,024 of 1,024 | +4.39 | 36 |
Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and event rates agree to 0.01 MH/s; no spills, no stack, temperature 58 to 64 C, no throttling. The verifier totals in the box's logs (about 10 ms per unit) are a naive scalar interpreter with a lazy mh_word per load on a shared Zen 4 core at load 13 and are not comparable to the project's Rust verifier (2.06 ms per unit on an M5 Max core, interleaved); the block delta is the number: about 1.1 microseconds per mm8 step per unit, 32 lanes x 32 byte products in plain loops. (At R = 512 the card issues about 258 G tiles per second, 2.6 x 10^14 multiply-adds per second, in the region of 40 percent of the 4090's dense int8 tensor peak, approximate from the published TOPS figure, so R = 512 is near the top of the free band, not its middle.)
What the rows say:
- The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (
latency-shadow-2026-10-06.md5). - The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op,
latency-shadow-2026-10-06.md5, item 4); the tensor block at its free-band ceiling costs a third of that. - That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at
k = 1the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (k_mmanear 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force. - The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (
family-probe.cumm8warp_ref) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that anmm8family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
Consequences per tier (B at R = 128, the largest R whose scalar verifier delta stays near 1 ms):
| Tier | Meaning | What to do |
|---|---|---|
| NVIDIA Turing and later (RTX 20 to 50 series), any memory size | rate unchanged, +11.5 W on a 4090 (+5.7 percent), per-pound unchanged, per-watt 0.95x | nothing; the block would not be adopted on this evidence |
| NVIDIA Pascal (GTX 10 series) | emulation at about 16 steps per mm8 (approximate): 1,024 per hash is about 16,000 counted steps, inside the latency shadow of a 1080-class card by the counted-ops rule, approximate; unmeasured |
the emulation path exists in the kernel text (IGNEUM_MM8_REF) and is bit-exact; a GTX 1080 row is owed if B ever proceeds |
| AMD RDNA 3 and 4 | the WMMA layout is unverified (status item 6); emulation otherwise at about 12 to 16 steps per mm8 |
the gate of section 7 rank 2 stays open; not worth running unless B proceeds |
| Apple | emulation at about 10 steps per mm8 (approximate), 10,000 counted steps at R = 128 inside the M5 Max's 130,000 ceiling |
owed; not worth running unless B proceeds |
| Rig, pool user | nothing changes in shares; a rig pays the block's watts per card | nothing |
| Node operator (verifier) | +1.14 ms per unit scalar at R = 128 (+55 percent of 2.06 ms), +4.39 ms at R = 512; a SIMD byte-dot path would cut it 4x to 16x (approximate) | the verifier cost per joule forced is the reason the scheme loses to the ALU shadow |
| A chip | must carry a tensor array per 32 lanes in flight; at the measured block energy it pays at most 0.23 microjoules per hash more than today at k = 1, 0.07 at k = 0.3 |
the chip's edge barely moves (5.3) |
5.2 state-dataset (scheme C), box 2 and igneum-build-1
Prototype: proto-newpow/state-dataset/ (README with every command and log; kernel_sd.cu = the pack's kernel plus igneum_leaves and igneum_build_sd, the hash kernel byte for byte the pack's; verify_sd.c the 32-lane C interpreter and the item check; cpu_rows.c the igneum-build-1 rows). Box 2: RTX 4090 24 GB, driver 595.91, nvcc 12.8, host EPYC 7352; igneum-build-1 EPYC 9454P at load 19 to 32 (shared), one core pinned, nice 19. Run 19:34 to 19:50 UTC. The leaf stand-in: one ChaCha12 block of (sigma, S, t, 0, tag) with S = K XOR 0x5a5a5a5a (the README defines it).
| Row | Control (mx8-genesis) | sd1 | Reading |
|---|---|---|---|
| Hash rate, 10 batches of 2^24, two passes | 63.083, 63.083 MH/s | 63.088, 63.087 MH/s | equal within 0.01 percent: the kernel is the same binary over a dataset of the same shape |
| Watts, mean after 10 s of a 20-s window; SM MHz | 207.6, 205.2 W; 2,745 | 207.9, 206.5 W; 2,745 | equal within the 2 W run-to-run noise; 3.27 microjoules per hash on this 4090 (the 5090 is 2.34 to 2.65) |
| MH/s per W | 0.304, 0.307 | 0.303, 0.305 | unchanged |
| Cache fill, GPU | 1.81, 1.85 ms | 1.85, 1.85 ms | unchanged |
| Leaf array (2^24 ChaCha12 blocks, 1 GiB) on the GPU | none | 5.49 ms | the synthetic stand-in; a real snapshot comes from the node |
| Dataset build, second pass | 30.55 ms | 31.98 ms | +1.43 ms (+4.7 percent): one coalesced 64 B read per item |
| Chunked build (leaves streamed from pinned host memory in 64 MiB or 256 MiB chunks, one stream) | 75.8 ms, 75.5 ms | PCIe-bound: 62 ms of copy (17.2 GB/s on this box) plus the build, partly serialised; two streams would hide most of the build (not done) | |
| Device memory during the build (context 395 MiB included) | 1,675 MiB | 2,699 MiB resident leaves; 1,741 MiB chunked at 64 MiB; 1,933 at 256 MiB | while hashing 1,803 MiB in every mode (leaves freed) |
| 2^24 fingerprint at base 0 | 7c28cfb06c5c65a9 (the pack's, both passes) | d5b0c16390cad0e8 (four runs) | |
| Pack vectors (3 warps, standalone and in batch); dataset head, word [MASK], 64 Mac samples | PASS | n/a (new values; sd1 base 0 lane 0 = b600edbed969becc) | the harness is the pack's |
| Build bit-exact: 1,024 random items, GPU words against the plain C host derivation | 1,024 of 1,024 | 1,024 of 1,024 (also after both chunked rebuilds); 64 device leaves = host leaves | |
Hash bit-exact: 4 dumped warps (128 lanes) against the C interpreter with lazy mh_word / mh_word_sd |
128 of 128 (and 32 of 32 against the Mac vector at base 0) | 128 of 128 | |
| Host cache FNV-1a 64 | 48c4f5bf24166b2e = the Mac's | same |
The CPU rows (igneum-build-1, ms per unit of 4,096 items, 100 units; two readings at load 19 and 31):
| Row | Reading 1 | Reading 2 | Per item |
|---|---|---|---|
| (i) 4,096 random 64 B reads from a 2 GiB resident leaf array | 0.108 ms (0.092 to 0.172) | 0.163 ms | 26 to 40 ns |
| (i) the same from an 8 GiB array (the year-12 size) | 0.139 ms | 0.209 ms | 34 to 51 ns |
| (ii) 4,096 leaves derived on the fly (one ChaCha12 block each) | 0.284 ms | 0.285 ms | 69 ns |
(iii) 4,096 x mh_item on the host cache, naive, no interleaving |
9.01 ms | 11.24 ms | 2.2 to 2.7 microseconds |
4,096 x (leaf read + mh_item_sd), naive |
9.31 ms | 11.64 ms | |
| Daily leaf array on CPU (2^24 blocks): one core / 32 threads | 1.67 s / 0.073 s | 1.68 s / 0.083 s | 100 ns per leaf |
| Host cache fill (256 MiB), one core | 0.447 s | 0.452 s | |
| Light-client sample: distinct items touched by lane 0's 128 loads | 128 of 2^24 | 128 openings x 25 levels x 32 B + 128 x 64 B = 108 KiB per lane; 3.4 MiB per 32-lane unit before dedup (arithmetic) |
Reading, against the gate: the project's Rust verifier does the 4,096 derivations in 2.06 ms on an M5 Max core by interleaving the eight dependent cache misses across items (the naive C figure here, 9 to 11 ms, is an upper bound and is not the production verifier); sd1 adds 0.11 to 0.21 ms per unit when the verifier holds the state in RAM (+5 to +10 percent of 2.06 ms; about +0.3 to +0.5 ms on a 2019-class core by the 2.5x rule), so the class v3 verifier plus sd1 sits at about 2.2 ms on the M5 Max core and about 5.7 ms on the 2019-class core, inside the 10 ms gate with the x8 margin intact. The snapshot's own cost is the state read on the node, not the leaf hashing (1.7 s on one core for 2^24 leaves).
Consequences per tier (sd1):
| Tier | Meaning | What to do |
|---|---|---|
| Home miner, one 8 GB card | hash rate, watts and MH/s per W unchanged (measured equal); the build must stream the leaves (resident leaves at the designed 2 GiB dataset would peak at about 4.6 GiB, approximate; chunked 64 MiB about 2.7 GiB, approximate, measured 1,741 MiB at 1 GiB), so the 8 GB mine-and-prove profiles of prover-tiers-real-cards.md keep their memory |
ship chunked only; the kernel already takes an item range, so chunking needed no kernel change |
| 12, 16, 24, 32 GB cards | unchanged in rate and watts; the daily build +1.4 ms resident or about 76 ms streamed (about 150 ms at 2 GiB, approximate) | nothing |
| AMD and Apple | the hash kernel is untouched, so the measured class v3 rates stand (18.9 MH/s on the 9070 XT, 27.1 on the M5 Max); the build kernels gain one 64 B read per item in OpenCL and Metal | the two build kernels in the emitter's other dialects |
| Rig | per card as above; one 2 GiB snapshot per day per rig, shared by its cards | the node beside the rig serves it over loopback |
| Pool user | a 2 GiB download per day from the pool, or the pool ships the built dataset; today a pool user needs nothing but the day key | the pool protocol (spec 09) gains a daily snapshot or dataset fetch |
| Solo miner | must run a node with execution (it already does for templates) and hold the state | nothing new beyond RAM |
| Node operator | +2 GiB RAM for the snapshot (+4 GiB at year 4), a daily serialisation pass, the verifier +0.11 to 0.21 ms per unit | the node RAM budget stays inside a 16 GB home node at launch sizes |
| Light client, holder | a verifiable sample of 128 state leaves per lane per block exists; checking it costs 108 KiB per lane fetched on demand, so it is an availability spot check, not a light verification of PoW | rank 6 of section 7 |
| Prover, rollup customer | nothing |
5.3 The chip rows on the measured numbers
Method: sim/horizon/new-pow/chip_rows.py (the chip-model-v3 section 5.4 f = 1 chip plus the block's measured energy times k; the card figures are 5.1's; every chip figure is arithmetic on chip-model-v3.md's cited and approximate memory figures). The denominator here is the 4090 measured tonight (3.19 microjoules per hash at the control, a card whose memory system is half the 5090's), not the 5090, so the R = 0 column reads 6.9x where chip-model-v3.md reads 5.1x for the 5090; the movement with R is the row that matters.
| Scheme and setting | Honest card microjoules per hash | Chip, GDDR7, k = 1 / 0.5 / 0.3 (microjoules) | Edge over the card, GDDR7 | Chip, one HBM3 stack, k = 1 / 0.5 / 0.3 | Edge, HBM3 |
|---|---|---|---|---|---|
| Class v3 today (R = 0, the 4090 control) | 3.19 | 0.466 | 6.9x | 0.321 | 9.9x |
| B at R = 128 (1,024 tiles per hash, block 0.18 microjoules) | 3.37 | 0.646 / 0.556 / 0.520 | 5.2x / 6.1x / 6.5x | 0.501 / 0.411 / 0.375 | 6.7x / 8.2x / 9.0x |
| B at R = 512 (4,096 tiles per hash, block 0.23 microjoules, the free band's top) | 3.42 | 0.696 / 0.581 / 0.535 | 4.9x / 5.9x / 6.4x | 0.551 / 0.436 / 0.390 | 6.2x / 7.8x / 8.8x |
For comparison, the ALU shadow at N = 100,000 on the 5090 (latency-shadow-2026-10-06.md 6, measured 3.27 microjoules at the 431 W cap) |
3.27 | 1.59 / 1.03 / 0.80 | 2.1x / 3.2x / 4.1x | 1.44 / 0.88 / 0.66 | 2.3x / 3.7x / 5.0x |
| C (sd1) at any size | unchanged (measured equal) | 0.466 | unchanged | 0.321 | unchanged; the f = 0 recompute chip must add 2 GiB of DRAM and becomes this chip |
Reading: B moves the f = 1 chip's edge by 1.4x to 2x at k = 1 and by 1.1x at k = 0.3, at a verifier cost of 1.1 to 4.4 ms per unit (scalar); the ALU shadow moves it by 2.7x at k = 1 and 1.4x at k = 0.3 at 0.17 ms per unit. On every axis the measured tensor block is the weaker lever, for the reason stated in 5.1 item 3. C moves nothing and claims nothing against the chip; its merit is elsewhere.
6. Verdicts
| Scheme | Verdict | Why, in one line |
|---|---|---|
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the mm8 family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
|
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
A, in full. The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.
NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the mm8 family as reserve R8 with the two-output correction, for datapath diversity, not for joulesC_PROSE
7. Ranked next steps
Hours are agent hours (the project lead's rule: Claude-side work takes hours). Each gate is a measurable pass line. Consequence per tier is the row's own. Ranks 2, 3 and 4 were written as B's gates before the ladder landed; after section 5 they are WITHDRAWN (B is not carried forward as class content; the rows stay so the reasoning is visible) and the live order is 1, 5, 6, 7, 8.
| Rank | Proposal | Evidence | Model | Hours | Consequence per tier | Gate |
|---|---|---|---|---|---|---|
| 1 | Scheme C as class v5 content: the daily dataset derives from the state snapshot (sd1), spec text for 1.8.5 (the leaf line), 1.12 (the snapshot's checkpoint and lead), 10 (the pruning-proof witness), 4.3's pause rule applied to the day boundary |
section 5.2: the hash kernel and rate are unchanged by construction, the build and verifier costs are the measured rows; every miner operation must hold state | the verifier row: 2.06 ms + the leaf-read row per unit; node RAM + the dataset size | spec 3 h; emitter and memhard.rs leaf line 2 h; node snapshot builder (canonical serialisation, one pass per day, the prepare hand-over) 6 h; fast-time harness across a day boundary and a finality pause 3 h; Devnet 2 crossing 2 h |
8 to 32 GB cards: no hash-rate change, +2 GiB device memory during the build only (streamable); rig: the same per card; pool user: a 2 GiB daily download or nothing; solo miner: a full node's state; node operator: +2 GiB RAM (+4 at year 4) and a daily leaf pass; holder: a daily random sample of state per block; prover, rollup customer: nothing | the fast-time 3-node network crosses a day boundary with a non-trivial state and no fork; verifier under 10 ms on a 2019-class core with the snapshot in RAM; Devnet 2 PASS over a day boundary |
| 2 | Scheme B's AMD gate: the WMMA iu8 fragment layout on the RX 9070 XT against the spec reference (the two-output tile), a PC 1 job with the card alone | status item 6 (layout unverified); section 5.1's bit-exact rows on NVIDIA | the spec layout as the definition; a lane permutation per vendor | 3 h (the probe exists: family-probe.cu mm8 row; the OpenCL twin needs the builtin path and the permutation) |
AMD 16 GB tier: decides whether RDNA 3 and 4 run B natively or emulate at about 12 to 16 steps per mm8 | 1,024 random units bit-exact on the 9070 XT, both tile outputs |
| 3 | Scheme B's Apple row: the emulated mm8 block in Metal on the M5 Max at R = 32, 128, 512, hash rate and watts under with-lock.sh measure |
section 3.2's vendor table: Apple is the honest card's worst case; the M5 Max binds at about 130,000 counted ops | 4 shuffles plus 4 dot4 emulations per step, about 10 steps per mm8 (approximate) | 3 h | Apple tier: the 5 percent rule decides the largest R the class can carry | rate within 5 percent of mx8 at the chosen R; bit-exact against the C reference |
| 4 | Scheme B as class v5 content beside C (mx8+sd1+mm8xR) once ranks 2 and 3 pass: the emitter's three dialects (inline PTX with the IGNEUM_MM8_REF fallback for sm_61 to sm_72, the AMD builtin behind a feature test, the Metal emulation), the verifier step in verify.rs, the stats and fuzz runs of TESTS.md on the class, the v5 switch beside v4 |
section 5.1 and 5.3: the block's measured watts and the chip rows at k | the chip rows of 5.3 | emitter 4 h; verifier 1 h; suites 2 h; node seam 2 h; Devnet 2 2 h | NVIDIA Turing and later, AMD RDNA 3 and later: native, the measured watts; Pascal, RDNA 2, Apple: emulation at the measured 5 percent check; pool user: nothing; chip: a tensor array per 32 lanes in flight | the six gates of counter-asic-2-rollout.md 7 plus the AMD and Apple rows; verifier under 10 ms on a 2019-class core at the chosen R |
| 5 | The reserve text correction: R8 mm8 consumes both tile outputs (two destinations) so a chip cannot halve the work |
section 3.2 (the half-tile shortcut) | arithmetic: 32 of 64 products consumed in the single-output form | 1 h (spec 1.13.2 text and the edge vectors) | none until era 4 or a signal | the R8 edge vectors re-cut for two outputs, bit-exact on the three vendors |
| 6 | The light-client state sample: a node RPC that returns, for a block, the 128 leaves of lane n with openings against R_d, and a client check | section 3.3 (the sample is real but modest: availability, not correctness) | 128 x 25 x 32 B = 102 KB per block fetched on demand (arithmetic) | 4 h | holder and light client: a spot check of state availability per block; node: one RPC | 128 openings verify against R_d on the fast-time network |
| 7 | The 2019-class core measurement (O-1.14) for the verifier under C and under B at the chosen R | every verifier row here is M5 Max or Zen 4 plus the 2.5x rule | the 2.5x rule stands in | 1 h once a core is found (a 2019 laptop or a rented older CPU box) | every tier: the gate that fixes R and the snapshot read budget | under 10 ms per unit, worst cold |
| 8 | Do not build scheme A (mining is proving) in any form; record the sampleability bound (proving work per segment over network hashes per segment: 8 percent at 1 GH/s, 0.08 percent at 100 GH/s) in the ledger beside F13 | section 3.1 and 4.1 | the bound's arithmetic | 0.5 h (a ledger row) | none | none |
One paragraph each.
1. Scheme C first. It is the cheapest change with a new property: the kernel text, the hash rate and the chip rows are untouched, so every measured number of Counter ASIC 2.0 and 3.0 stands; what changes is the daily build and the node. The cost is a canonical state serialisation in consensus, which is why the fast-time harness must cross a day boundary with state and a finality pause before the Devnet 2 crossing. The pool caveat is stated: the design forces the operation, not the card, to hold state.
2 and 3. B's two vendor gates. B is bit-exact tonight on NVIDIA (section 5.1) and in every emulation; it is not a class candidate until an AMD card reproduces the spec layout and the Apple emulation's hash-rate cost is measured. Both are short jobs with the card alone.
4. B beside C. The order matters: C lands first because it needs no vendor work; B joins the same class byte or the next one once ranks 2 and 3 pass, with R set from 5.1 and the 2019-core row.
5 to 8. Small, named, and each closes a thread this lane opened.
8. Open questions and what could not be run
| Question | Why it could not be closed tonight | What closes it |
|---|---|---|
| The 2019-class core (O-1.14) | no such core in the fleet; igneum-build-1 is Zen 4, the Mac is M5 Max; the 2.5x rule stands in | rank 7 |
| The AMD WMMA fragment layout | PC 1 is the project lead's desk and the AMD rows were owed all day (status file); no AMD card on RunPod or Vast tonight (fleet agent) | rank 2 |
| The Apple emulation of mm8 | a Metal emulation kernel is a 3-hour job and the Mac measure lock was free; not started because the AMD gate decides first whether B proceeds | rank 3 |
| The canonical state serialisation for C | a design item that touches the exec layer (igneum/exec), out of this lane's files |
rank 1 |
| The pruning-proof witness for C's historical PoW | spec 02 and 10 items | rank 1 |
| Whether the tensor path's marginal energy on the 5090 differs from the 4090's | one card measured (box 1); the 5090 is on the project lead's desk | a PC 2 job with the same run.sh |
| The Ampere row (3060, 3080, 3090) | every Ampere card of the fleet was mining and proving the live devnet; a loaded 3090 was offered and declined (a loaded card's rate is not a number) | one quiet Ampere pod |
| The verifier with a byte-dot instruction (VNNI, NEON udot) | the C reference is scalar | 1 h: an AVX-VNNI and a NEON path in verify_ref.c, measured on both cores |
| Scheme C's leaf array on an 8 GB card at the 2 GiB design size | the build holds dataset + cache + leaves on the device (4.3 GiB) unless chunked | the chunked build (section 5.2 says whether it is trivial) |
9. Summary for the coordinator
This lane wrote three candidate proofs of work for Igneum, reviewed each in two personas, prototyped the two that survived on real cards, and measured. (1) Scheme C, the dataset derived from the chain's execution state, is the one to carry forward as the class v5 candidate: measured on an RTX 4090 the hash rate and watts are unchanged to the third decimal (63.08 MH/s, 207 W, the kernel is byte for byte the shipped one), the daily build grows by 1.4 ms resident or about 76 ms streamed, the verifier by 0.11 to 0.21 ms per unit with the state in RAM, every GPU item and lane agrees with the plain C derivation (1,024 of 1,024, 128 of 128); it forces every mining operation to hold state, names 128 random state leaves per lane per block, and removes the f = 0 recompute chip as a category; its costs are a 2 GiB daily delivery for pool miners and three spec items (canonical serialisation, a pruning-proof witness, the pause rule). (2) Scheme B, a tensor-shaped shadow of int8 8x8x16 tiles, is bit-exact on NVIDIA at every rung (2^24 lanes PTX = reference, 1,024 lanes CPU = GPU, R = 8 to 512) and free in hash rate to 4,096 tiles per hash, and that is exactly why it fails as an anti-chip lever: it costs the honest card 0.056 to 0.70 pJ per multiply-add and at most 0.23 microjoules per hash, so the f = 1 chip's edge moves from 6.9x to 4.9x at k = 1 against the 4090 where the class v4 ALU shadow moved the 5090's from 5.6x to 2.1x, at 26x the verifier cost per joule forced; never as class content, kept as the evidence and the two-output correction for reserve family R8. (3) Scheme A, mining is proving, is never: one proof per segment is not a distribution of puzzles; the useful fraction is bounded by proving work per segment over network hashes per segment (8 percent at 1 GH/s, 0.08 percent at 100 GH/s on this chain's measured figures); every form that is not "hold the trace" dies on 2.9 MB of openings per block or a 32 to 40 ms proof verify, and "hold the trace" is scheme C with a worse data source.
- Scheme C costs nothing in hash rate or watts (63.083 against 63.088 MH/s, 205 to 208 W on the 4090, measured) and 0.11 to 0.21 ms per unit of verifier time (measured on igneum-build-1), so class v5 can carry it; the model is
item_sd(t) = item(t) with s ^= leaf(t)over the day's certified state root. - The tensor shadow moves the f = 1 chip's per-joule edge by at most 1.4x (6.9x to 4.9x at k = 1, R = 512) for 4.39 ms of scalar verifier per unit, against 2.7x for 0.17 ms from the ALU shadow; the model is
chip_rows.pyon the measured block energy of 0.23 microjoules per hash. - Mining cannot be proving on a lottery with a 10 ms verifier: the useful fraction is bounded by 5 GPU-seconds of proving per 8-second segment over the network's hashes in that segment, 8 percent at 1 GH/s and falling with the hash rate; the model is that ratio on
prover-tiers-real-cards.mdand bench-log 2582.