The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh). The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge. Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
40 KiB
Counter ASIC 3.0 item 2: a random item-derivation program per day
6 October 2026. Worker derive (branch ca3-derive), under the brief of docs/plans/counter-asic-3.md item 2 and
the history audit's addition 2 (docs/analysis/asic-resistance-history.md section 4.3). Everything here is
PROPOSED: a prototype behind a load class (dr736), measured on the Mac and on PC 2, written as a reserve entry for
spec 1.13.2 (section 6) that lives in this file until the founder's word. Nothing is published and no vector of class v2
or v3 moves (section 3.4).
What it does, in one line: the fixed-shape mixer M_r of spec 1.8.4, applied 72 times per item under class v3,
is replaced by nine straight-line programs of 736 instructions drawn once a day from the day key stream, so the chip
that holds the cache on die must run an arbitrary program instead of a wired pipeline; the 8 dependent cache reads
per item, the cache, the loads and the hash kernel are untouched.
0. The verdict first
| Gate | Number | Bar | Result |
|---|---|---|---|
Verifier per 32-lane unit, one M5 Max core, with-lock.sh measure, load average 4.9 / 4.5 / 5.3 |
4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class dr368 (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
| Bit-exact: Metal, Apple OpenCL and CUDA (NVRTC, PC 2) against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples, 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, all three compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal and CUDA) | equal | passes on three compilers (5.2, 5.4a) |
| Daily 1 GiB build, M5 Max, Metal, measure lock; RTX 5090, PC 2 (the card still mining, 5.4a) | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%); 5090 42 / 32 ms against x8's 40 and v2's 46 on the loaded card | under 1 s on every discrete card | passes on both; the 9070 XT OWED (PC 1 is the founder's desk today) |
| NVRTC compile per pack, RTX 5090 | 1,266 ms against x8's 164 (+1.1 s): the item function is inside every hash-kernel and race-variant compile | the 38 s compile-ahead budget at the 600-s epoch floor | the one-module-per-day fix is a requirement of the class (5.4a) |
| Hash rate, M5 Max, Metal, measure lock; RTX 5090 (loaded) | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16; 5090 61.1 / 62.1 against x8 62.0 and v2 62.3 | equal within noise | equal (0.3% Mac): the hash kernel does not change |
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
core measurement (O-1.14) lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core
(pass) against about 12 ms on the approximate laptop row (fail). The class that passes both rows today is dr368
(2.69 ms), at the x4-equivalent op count, with the chip row at 0.58x bare. The way to the 736 figure under the gate
on a laptop is the JIT (section 4.3), which is out of scope and named with its risk.
The verifier headroom each class leaves under the 10 ms gate, the budget other 3.0 items (item 8, program work in the latency shadow: 32 N CPU ops per unit at N ops per hash) may spend on top of this one:
| Class | ms per unit, this M5 Max core (steady / worst cold) | Headroom under 10 ms (steady / worst cold) | On a 2019-class core (2.5x, approximate) | Headroom there (steady / worst cold) |
|---|---|---|---|---|
| x8 (class v3 today) | 2.06 / 2.18 | 7.9 / 7.8 ms | 5.2 / 5.4 | 4.8 / 4.6 ms |
| dr736 | 4.88 / 5.24 | 5.1 / 4.8 ms | 12.2 / 13.1 | none: over by 2.2 / 3.1 ms |
| dr368 | 2.69 / 2.90 | 7.3 / 7.1 ms | 6.7 / 7.2 | 3.3 / 2.8 ms |
So item 8's N is bounded by 4.8 ms on this core and by nothing on the approximate laptop row under dr736, by 7.1 and 2.8 ms under dr368, by 7.8 and 4.6 ms under x8 alone; the two items share one budget and the laptop row is the one that binds, which is one more reason the O-1.14 measurement comes before either goes genesis-live.
1. Why: what the chip model says the fixed shape is worth
docs/analysis/chip-model-v3.md section 2 prices the on-die-cache recompute chip at 50 T op/s: class v3 (x8) costs
it 1,198,080 integer ops per hash (72 mixer applications x 128 items x about 130 ops), 41.7 MH/s, 0.31x the 5090's
136.1 MH/s bare, and 0.92x with "the 3x fixed-function factor (approximate, from memory: 2x to 5x is the usual
credit for a pipeline with no scheduling or divergence)". That credit is the mixer's fixed shape: the chip unrolls
the 72 applications into a wired pipeline with the day's constants baked in, no instruction fetch, no register
file, no operand muxes. RandomX's answer (vendor/RandomX/src/superscalar.cpp is not in this tree; read on
6 October 2026 from the upstream repository at commit 7607fb2 into the session scratchpad; the design argument is
doc/design.md, history section 2.4) is SuperscalarHash: the dataset item derivation is itself a random program
drawn from the cache key, 8 programs of about 450 instructions per item, so a light-mode chip "becomes a CPU".
RandomX's generator (generateSuperscalar, lines 653 to 850) schedules for a superscalar x86 core: it picks a
decode-buffer configuration per cycle, selects a source register that is ready at the cycle and a destination
that is not the source and was not last written by the same op group (selectDestination, line 495: no
"xor r,r2; xor r,r2", no "ror r,C1; ror r,C2", no two multiplies in a row on one register), and then computes the
program's ASIC latency as the longest dependency chain (lines 810 to 824) and sets the address register to the
register with the highest one. The item init is rl[0] = (itemNumber + 1) * superscalarMul0, the other seven
registers rl[0] ^ superscalarAdd_i, then per cache access: run the program, XOR the mix block in, next address
from the address register (dataset.cpp initDatasetItem, lines 164 to 190).
What carries over here and what does not: the per-day program, the acceptance by construction, the dependency chain and the address register idea carry over; the x86 port scheduling does not (our verifier interprets the program for 32 items at once and our miners compile it for a GPU, so the schedule that matters is the GPU's), and the latency bound RandomX relies on (the program's critical path against DRAM) is replaced by ours: the 8 dependent cache reads per item, which are untouched.
2. The design
2.1 The draw
One SplitMix64 stream seeded with K[0] | (K[1] << 32) (the day key, spec 1.8.1), the 40 draws of 1.8.4
(ROT, MUL, RC) first, exactly as today (the item init s[8 + i] = t * MUL[i] + RC[i] still uses them), then
the program: DERIVE_PROGRAMS = 9 round programs (one before each of the 8 cache reads, one after the last) of
derive_len instructions each, four draws per instruction in a fixed order:
| Draw | Range | Sets |
|---|---|---|
below(100) |
op roll | the form, by the weight table of 2.2 (cumulative) |
below(15) |
destination roll | d: the 15 registers other than the chain c, in ascending order (roll >= c adds one) |
below(14) |
third-register roll | b: the 14 registers other than d and c, ascending (two skips); used by andx and orx, consumed by every form |
next() |
the immediate | k = 1 + (low32 mod 31) for the rotate forms; imm = low32 for addc, xorc; imm = low32 OR 1 for mulc, mulc2; consumed by every form |
The chain register c is s[0] (the address word) at the start of each round program and the destination of the
previous instruction after it. Every instruction reads c and writes d != c; the next instruction's chain is
d. So no two instructions of a program can run in parallel (the "every instruction consumes the newest result"
rule of the brief, SuperscalarHash's chain made strict), and no two consecutive instructions write one register,
which is what removes the mergeable pairs SuperscalarHash's selectDestination guards against. Four draws per
instruction whether the form uses them or not, so the stream position of every draw is fixed by its index and a
future change to one form's draw changes no other draw (the rule 1.13.1 follows for epoch_len). A rejected
candidate (2.4) is followed by the next: the stream continues, attempt + 1, as the program generator of 1.4.6.
Source: igneum-pow/src/derive.rs (DeriveProgram::draw_candidate, draw), memhard.rs (MixParams::with_shape).
2.2 The instruction set: twelve two-register forms, fixed at genesis
c the chain, d the destination, b the third register, k in 1..31, i a 32-bit constant (odd for the
multiplies). All arithmetic modulo 2^32, no division, no float, no data-dependent branch. Every form is a bijection
on the 16-word state (the old d enters through +=, -=, ^=, an odd multiply, or a rotation of itself; c
and b are not written), so a program loses no entropy, the property M_r has.
| Form | Semantics | Weight (percent) | GPU ops | Chip ops | Multiply |
|---|---|---|---|---|---|
add |
d += c |
14 | 1 | 1 | |
sub |
d -= c |
10 | 1 | 1 | |
xor |
d ^= c |
14 | 1 | 1 | |
mul |
d *= (c OR 1) |
10 | 2 | 1 (the OR is a wire) | yes |
rot |
d = rotl(d, k) + c |
10 | 2 | 2 | |
xrot |
d = rotl(d ^ c, k) |
10 | 2 | 2 | |
addc |
d += c + i |
6 | 2 | 2 | |
xorc |
d ^= c ^ i |
6 | 2 | 2 | |
mulc |
d = (d ^ c) * i (the per-word form of M_r, the chain in place of the round constant) |
8 | 2 | 2 | yes |
mulc2 |
d = d * i + c |
4 | 2 | 2 | yes |
andx |
d ^= (c AND b) |
4 | 2 | 2 | |
orx |
d += (c OR b) |
4 | 2 | 2 |
Mean 1.62 GPU ops and 1.52 chip ops per instruction, 22% multiplies. and and or enter only as andx and
orx (a destructive d &= c would lose bits; the XOR and add of a conjunction keep d invertible). The forms are
the brief's set (add, sub, mul-lo, xor, rotate by 1..31, and, or, the ARX-multiply forms of M_r); mulhi is in
the lottery hash's families and bit-exact on the three vendors, and is left out of the derivation on purpose so
every form is one that the three compilers lower to a single integer instruction (section 3.3).
2.3 Length and the floors: the x8-equivalent op count
The coordinator's rule for item 2: the total op count per item equals or exceeds today's x8 count, so the chip's
budget row does not fall. The x8 mixer counted from the code (memhard::mixer): 16 x (xor, add, mul) + 8 quarter
rounds x 12 = 144 ops as written, 128 with the RC[i] + rk adds hoisted as constants (the chip and every compiler
do that), 16 multiplies; 72 applications per item = 10,368 ops as written, 9,216 hoisted, 1,152 multiplies
(chip-model-v3.md section 1 prices 130 per application from the spec text; the item 3 worker counted the same
144 / 128, coordinator's note of 6 October). The floors are those three. derive_len = 736 gives 6,624
instructions per item, expected 10,731 GPU ops, 10,068 chip ops and 1,457 multiplies (the genesis day draws
10,659 / 9,992 / 1,461; the devnet day 10,701 / 10,083 / 1,362), 7 to 22 standard deviations above the floors, so a
rejection on a floor is a rare event and the test exists for the degenerate class. The class name is dr736; the
half-length dr368 (the x4-equivalent) is measured beside it as the fallback with its floors scaled.
2.4 The acceptance test
DeriveProgram::check, on every candidate; a rejection draws the next attempt from the stream:
| Test | Rejects | Expected rate at 736 |
|---|---|---|
| every register written in every round program | a register no instruction of a round program writes (its init word would never enter the chain within that round) | 16 x 9 x (14/15)^736 = under 10^-20 |
| at least 8 distinct rotation amounts across the item's programs | all rotations equal (the brief's degenerate draw; the mixer's "ROT draw of eight equal values is possible and untested", spec 1.8.4) | about 1,300 rotate forms over 31 values: never |
| chip ops >= 9,216, GPU ops >= 10,368, multiplies >= 1,152 per item (scaled to the length) | an op draw under the x8 count | 22, 9 and 9 standard deviations below the mean: never |
structural (asserted, hold by construction): d != c, b distinct from both, k in 1..31, odd multiplier constants |
a generator bug | n/a |
The generator panics after 64 rejected candidates in a row (MAX_ATTEMPTS), which the rates above put beyond
any day the chain will see; the panic is the right failure (a node that cannot derive the day's program cannot
verify, and must say so rather than guess). The weights are fixed at genesis; only the order, the registers and
the constants are drawn, so the family mix of a program cannot be steered by the draw.
2.5 The item, with the program in place
Spec 1.8.5 under the derivation class (derive_len nonzero, mixer_mult unused):
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
s = P_r(s) round program r, the chain starting at s[0]
a = s[0] AND (2^(C - 4) - 1) cache line index, as today
s[i] = s[i] XOR cache[line a][i] for i in 0..15
s = P_8(s)
item(t) = s
The 8 dependent cache reads per item are exactly today's: the address of read r is s[0] after program r, and
s[0] depends on every earlier read through the chain (every round program writes every register, 2.4, and the
XOR of the line into all 16 words feeds the next program). The verifier's latency part (8 dependent misses per
item, overlapped across the up to 32 items of a load, spec 1.11) is unchanged, which the measurement shows: the
difference against x8 is the ALU part only (section 5.1).
3. The prototype
3.1 Where it lives
| Item | Where |
|---|---|
DOp, DInstr, DeriveProgram (draw, check, counts, fingerprint), the SoA interpreter run_round (pair dispatch), the scalar reference run_round_scalar, the text forms instr_text and instr_line |
igneum-pow/src/derive.rs |
Shape::derive_len, Shape::is_derived, MixParams::derive (drawn after the 40 mixer draws), derive_items dispatching to derive_items_program |
igneum-pow/src/memhard.rs |
LoadClass::derive_len, LoadClass::DR736, with_derive, parse and name dr<len>, the program id (derive/ + the length) |
igneum-pow/src/generator.rs |
mh_round_0..8 and the program-driven mh_item in memhard.h, memhard.metal and kernel.cl; IGNEUM_DERIVE_* in program.h; derive_len, the op mix, the floors' counts and the nine programs (one line per instruction) in program.json |
igneum-pow/src/emit.rs |
--class dr736 (or any dr<len>) on every command; the bench prints the program's counts |
igneum-pow/src/main.rs |
examples/derive_perf.rs: the interpreter's cost per instruction per batch, drawn program against uniform programs |
igneum-pow/examples/ |
Packs dr736-genesis (seed igneum-genesis, day 2026-10-03, program id 72c1d8048aef9542, program fingerprint 463535d01511350d) and dr736-devnet-epoch0 (the devnet epoch 0 and day seeds, 7f4a5ca0a3637820, 771868df4e64d6ab); generator 2 with the class in the id, the x4-record shape, not a class v3 pack |
proto-cuda/packs-ca3-derive/ |
Tests: src/derive.rs (5), tests/derive.rs (7: by hand on a small cache, batches, v2 and v3 untouched, the stream and the class, determinism and the pack text, stats beside x8, the text forms against the scalar reference, the word path) |
igneum-pow |
| The PC 2 job | relay/playbooks/ca3-derive-pc2.ps1 (section 5.4) |
3.2 The verifier without a JIT: the word-major interpreter
The verifier derives up to 32 distinct items per load (one per lane of the unit, MemhardCpu::fetch). The
interpreter keeps the 32 item states word-major (st[reg][lane], 2 KiB) and runs each instruction across the
whole batch in one straight loop the compiler vectorises (NEON add.4s, mul.4s, ushl.4s and so on: 1,180
such instructions in the example binary), so the dispatch is paid once per instruction per batch, not per item;
the cache reads of the batch are issued together after each round program, as the fixed-mixer loop does, so the 8
dependent misses of independent items overlap. The dispatch is on PAIRS of instructions (144 arms, one indirect
branch per two instructions): a drawn op sequence is random, the predictor misses most dispatches, and pairing
halves the misses per instruction. Measured with examples/derive_perf.rs (a functional run, load average 3.9 to
4.9): 10.7 ns per instruction per batch cold, 7.18 warm with single dispatch, 4.98 with pair dispatch; a uniform
program of one form (predictable dispatch) 3.4 to 4.3 ns, so the body is about 3.5 ns and the remaining dispatch
cost about 1.5 ns. 6,624 x 4.98 ns x 128 batches = 4.2 ms per unit of interpreter time; the measured 4.88 ms
includes the latency part and the transposes.
3.3 Bit-exactness on three vendors, by construction
Each form is one C statement on uint with +, -, ^, *, |, & and the memhard core's mh_rotl (a
shift pair, n in 1..31 at every call site), the same text in Metal, CUDA C and OpenCL C, every operand a 32-bit
unsigned integer: the same argument as spec 1.14 for the lottery hash's families, which have run bit-exact on the
three vendors since 4 October. The program is emitted as nine functions of straight-line statements (196 KB of
memhard.h per day); NVRTC, the Metal compiler and the OpenCL compilers see no loop, no branch and no call inside a
round program. Measured: section 5.2 (Metal and Apple OpenCL), 5.4 (CUDA).
3.4 The v2 and v3 paths are untouched
Shape::derive_len is 0 and LoadClass::derive_len is 0 on V2, MX4, MX8 and V3_CLASS; MixParams::derive
is None; derive_items takes the fixed-mixer loop as before; the emitter's text for a shape without a program is
unchanged. cargo test -p igneum-pow (commit acb96ee): 58 lib, 7 derive, 4 mixer, 19 packs (every pinned pack of
v2 and v3 regenerated and compared byte for byte), 7 scratch, all green.
4. What the chip keeps, and the fallbacks
4.1 The fixed-function allowance after the change
The chip of chip-model-v3.md now has to execute, per item, 6,624 instructions from a 12-form set over a
16-entry register file, with three operand fields and a constant per instruction, in an order and with operands
that change every day. That is a sequencer: an instruction store (6,624 x 9 bytes = 60 KB per day program, in
SRAM beside the cache mirror), a register file with three read ports and one write port, a 32-bit ALU with a
multiplier and a barrel rotator, operand muxes, and a program counter. A GPU streaming multiprocessor is the same
machine with a wider register file and a warp scheduler. What the chip keeps over the GPU: no warp scheduler, no
operand collector, no instruction cache hierarchy for a 1 MB kernel (the GPU's day program is about 60,000 SASS
instructions per thread; whether the 5090's instruction cache holds it is in the build time of 5.4), no graphics
or float units idle on the die, and the day's constants folded into the instruction store. What it loses: the
wired pipeline (the mixer's 72 applications as 72 stages with the constants in the wires, no fetch, no register
file, no crossbar), which is the thing the 3x credit paid for. The chain rule adds a second loss: with no
intra-item parallelism a single engine finishes one instruction per cycle at best and its multiplies serialise
unless it interleaves items, which costs a register file per item in flight (RandomX's light-mode argument,
history 2.4: a chip paying "760 cycles and 1,240 multiplies per item"). ProgPoW claimed 1.1x to 1.2x for exactly
this kind of chip ("conventional compute chips gain little on ProgPoW", Bob Rao's hardware audit, history 2.4,
[S70]; EIP-1057's own claim 1.1x to 1.2x, [S67]). So the allowance this document carries is 1.2x (ProgPoW's
claimed range, cited) with 1.5x as the cautious upper bound (approximate, mine), against the 3x of the fixed shape
(approximate, from memory, M16). Section 7 prices all of 1.0x, 1.2x, 1.5x, 2x and 3x.
4.2 What it does not change
The partial-store chip (item 1): a chip that stores the dataset and never derives items pays nothing for the
program; item 1's rows stand on their own. The cryptanalysis question (item 3) changes shape: instead of one
fixed M_r to attack, the attacker gets a fresh random ARX program every day, which is RandomX's bet and is
untested here; the day's programs can be audited by the same tools as random ARX ciphers, and the acceptance test
is the place to add a structural rule if one is found. The era draws and the cache growth are untouched.
4.3 The JIT, named as the fallback with its risk
A per-day JIT for the verifier (emit NEON or AVX2 code for the 32-lane batch, the chain row kept in registers
across instructions since c is always the row just written) would remove the 1.5 ns dispatch and about a third
of the 3.5 ns body (the chain row's loads), about 2.5 ms per unit on this core, the x8 figure. Its risk: a code
generator in the consensus path on two architectures (arm64, x86-64) whose output must equal the interpreter's
bit for bit, writable-executable memory in a node and in every pool verifier, a new attack surface the history's
lesson 9 says to audit before launch, and a second implementation per platform to keep in lockstep. Out of scope
for this item; it is the route to 736 under the gate on a laptop core if the measurement of O-1.14 confirms the
approximate row, and dr368 is the route that needs no JIT.
5. Measurements
All under the locks of the brief; every row says its lock and the load average. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0. Commit acb96ee (the code and the packs), packbench built from this worktree.
5.1 The verifier per unit, one M5 Max core (with-lock.sh measure, one session, 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end)
Script measure-ca3-derive.sh (session scratchpad): igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <v2|mx8|dr736> --warps 50, two rounds, then the devnet seeds (--epoch-hex edc4fa84...fb07 --day-hex 69676e65756d2d6461792ffa50000000000000) and dr368 once.
| Class | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | Against x8 | Items per unit | Ops per item (GPU / chip / multiplies) |
|---|---|---|---|---|---|---|
| v2 (igneum-genesis) | 0.598 / 0.594 | 0.697 | 1 | 4,096 | 1,296 / 1,152 / 144 (9 x 144, hoisted 9 x 128) | |
| x8, mx8 (class v3) | 2.061 / 2.063 | 2.179 | 3.46x | 1 | 4,096 | 10,368 / 9,216 / 1,152 |
| dr736 | 4.875 / 4.944 | 5.241 | 8.2x | 2.37x | 4,096 | 10,659 / 9,992 / 1,461 |
| x8, the devnet seeds | 2.078 | 2.155 | 4,095 to 4,096 | |||
| dr736, the devnet seeds | 4.872 | 5.241 | 2.34x | 4,096 | 10,701 / 10,083 / 1,362 | |
| dr368 (the x4-equivalent fallback) | 2.692 | 2.898 | 4.5x | 1.31x | 4,096 | 5,350 / 5,004 / 752 |
Reading: the v2 row reads 0.59 to 0.60, the quiet readwidth night's 0.604 to 0.626 and 6.4a's 0.607 to 0.611, so this session is a quiet-core figure and no scaling applies. At the same chip-op count as x8 (plus 8%), the interpreter costs 2.4x what the compiled mixer costs: 4.98 ns per instruction per batch (3.2), of which about 1.5 ns is dispatch and 3.5 ns the vector body, against the mixer's compiled straight-line loop over the same batch. The latency part is the same in every row (the 8 dependent misses per item; the dr368 row at half the instructions saves 2.2 ms of the 4.9, which puts the latency-and-transpose share at about 0.5 ms). The 10 ms gate keeps 5.1 ms (worst cold 4.76 ms) at 736 on this core and 7.3 ms at 368.
5.1a Verification throughput per tier (the form of mixer-x4.md 6.5)
| Figure | x8 (this session) | dr736 | dr368 | Note |
|---|---|---|---|---|
| ms per unit, quiet M5 Max core (measured, this session) | 2.06 | 4.88 | 2.69 | |
| ms per unit, 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) | 5.2 | 12.2 | 6.7 | the figure that fixes the gate is a measurement, not this row |
| Shares per second per core (quiet M5 Max) | 485 | 205 | 372 | |
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second) | 4.5 | 10.7 | 5.9 | a pool verifying two units at once would halve the dispatch share (3.2); a 64-lane interpreter is a follow-up, unmeasured |
| Node: worst cold single unit (per block) | 2.2 ms | 5.2 ms | 2.9 ms | a block's verification stays under the 1 s block time by 190x |
| IBD over 108,000 headers on one core | 3.7 min | 8.8 min | 4.8 min | laptop (approximate): 9.4 / 22 / 12 min |
| Margin left under the 10 ms gate (worst cold, this core) | 7.8 ms | 4.8 ms | 7.1 ms | on the laptop row (approximate): 4.8 / none (over by 2.2) / 3.3 ms |
Consequences per tier: every miner tier is untouched by the verifier (the miner never runs it); a pool operator pays 2.4x the cores at 736 (11 cores for a 22,000-member pool against 4.5) or 1.3x at 368; a node on any 2026 core verifies a block in 5 ms; a node on a 2019-class laptop core is the open question (O-1.14), and at 736 the approximate row says it misses the gate, so 736 does not go genesis-live on an approximation, and 368 passes it.
5.2 Bit-exactness on the Mac (with-lock.sh run, 08:38 to 08:41 local)
packbench --pack <dir> --batches 1 --batch-log2 24 --group 256 (Metal, built from this worktree) and
igneum-bench-cl-dr736-genesis --bench-pack --pack <dir> --batches 1 --batch-log2 24 (Apple OpenCL, proto-opencl/ build.sh on the pack through a temporary link). Vectors are the Rust interpreter's.
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word [MASK], 64 samples | Vectors | Fingerprint 2^24 | Compile | 1 GiB build, GPU ms (run lock, indicative) |
|---|---|---|---|---|---|---|---|
| dr736-genesis | Metal | 48c4f5bf24166b2e PASS | head and last PASS (packbench checks no samples) | 3/3 standalone, 3/3 in batch | 50e3eaa779da4f1e | 784 ms (276 on the second process, 1 ms once the shader cache has it) | 38.2 / 38.3 |
| dr736-genesis | Apple OpenCL | PASS (head, last line, FNV) | head, word [268435455], 64 samples PASS | 96 of 96 lanes | 50e3eaa779da4f1e | (in the 385 ms prepare) | 57 ms wall |
| dr736-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 9553f6d5c667205a | 763 ms | 37.4 |
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the derived dataset (head, word [MASK], the 64 samples through OpenCL), on every vector lane and on the 2^24-output fingerprint of dr736-genesis across both compilers; the devnet-seed pack agrees on Metal. The Apple OpenCL compile of a 6,624-statement item function is inside a 385 ms prepare; the Metal compile is 0.75 to 0.8 s cold per day (the memhard library is the day's, compiled once a day, not per epoch) and 1 ms from the shader cache.
5.3 The daily build and the hash rate on the M5 Max (with-lock.sh measure, the session of 5.1, three rounds)
packbench --pack <dir> --batches 2 --batch-log2 22 --group 256, mx8-genesis then dr736-genesis, three rounds.
| Pack | Compile (round 1 / 2 / 3) | 1 GiB build, GPU ms (round 1 / 2 / 3) | MH/s GPU (round 1 / 2 / 3) | Vectors, self-tests |
|---|---|---|---|---|
| mx8-genesis (class v3, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | 3/3 + 3/3, PASS |
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | 3/3 + 3/3, PASS |
Reading: the build is 29 ms against 22 (+32%, +7 ms): the Mac's build was latency-bound at x1, x4 and x8 (21 to 22 ms at every multiplier, mixer-x4.md 6.4) and the serial chain per thread now shows (no intra-item ILP for the compiler to schedule), still 34x under the 1 s bar. The hash rate is the x8 rate within 0.3% (27.06 to 27.13 against 27.08 to 27.16), as it must be: the hash kernel only loads. Consequences per tier: a daily build of 29 ms on Apple silicon costs nothing to any tier; the 5090's figure is section 5.4, the 9070 XT's is OWED (PC 1 not released today; its x8 build was 72 to 77 ms and arithmetic-bound at none of x1, x4, x8, so the chain's cost there is the open number); the integrated tier is the one to watch (5.5).
5.4 RTX 5090, PC 2 (one job, relay/playbooks/ca3-derive-pc2.ps1)
PENDING at the time of writing: the job waits for /tmp/igneum-devnet/pc2-ca3.clear (the proving agent's
30-minute measurement) and the pc2-ca3.lock. The job downloads the packs zip itself (sha256
aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676: dr736-genesis, dr736-devnet-epoch0, mx8-genesis,
v2-genesis-mh), runs the installed app's igneum-worker-cuda.exe (NVRTC compiles each pack's own text) with the
NVIDIA card off in the app only under test (its key from settings.json), --check for the nvrtc .. cache .. dataset .. ms line and the self-test, --bench at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
The rows land in the bench-log entry when the closing report is read.
5.4a RTX 5090, PC 2: the rows (job run-ca3-derive-pc2-20261006, published 08:26:08Z, ran 08:26:43 to 08:28:49Z, done exit 0, 126 s; lock taken 08:25:55Z, released 08:29:08Z)
The installed 0.3.11 worker (sha256 2b3b8c92...c2674c, NVRTC 12.8, driver 13.3, sm_120, --bench present) on the
self-fetched zip (sha256 verified on the PC). THE CARD DID NOT COME OFF: the job read the card key from settings.json
as the brief asks (nvidia:0:NVIDIA GeForce RTX 5090), the mixer job of 5 October had used the state's key
(nvidia:NVIDIA GeForce RTX 5090, no index), and after the POST and 90 s one igneum-worker-cuda process was
still running (worker processes left 1; gpu-before 431 W, 3,896 MiB used, the prover's GPU server ON as the
clear file requires). So every row below was taken with the installed miner still on the card: the v2 control
reads 62.3 MH/s against its unloaded 136 to 137, and the same load sits under every pack, so the ratios between
the rows are the measurement and the absolute rates and build times are not (the key-format mismatch is filed in
section 8; the fix is to try both key forms and to confirm by the process list before measuring).
| Pack | NVRTC compile | Cache | 1 GiB build | Self-test | Fingerprint 2^24 (the Mac's) | MH/s, block-warps 1 / 8 (loaded) |
|---|---|---|---|---|---|---|
| v2-genesis-mh (control) | 167 ms | 6 ms | 46 ms | PASS (cache FNV 48c4f5bf24166b2e, head, word [MASK], 64 samples, 96 lanes) | 25f96e7dce90bd4e (equal) | 62.26 / 61.34 |
| mx8-genesis (x8, class v3) | 164 ms | 4 ms | 40 ms | PASS | 7c28cfb06c5c65a9 (equal) | 61.98 / 60.53 |
| dr736-genesis | 1,266 ms | 5 ms | 42 ms | PASS (cache FNV 48c4f5bf24166b2e, 64 samples, 96 lanes) | 50e3eaa779da4f1e (equal to Metal and Apple OpenCL) | 61.08 / 58.51 |
| dr736-devnet-epoch0 | 1,266 ms | 6 ms | 32 ms | PASS (cache FNV 448274a57f508cbc, 64 samples, 96 lanes) | 9553f6d5c667205a (equal to Metal) | 62.15 / 61.44 |
Reading. Bit-exactness on CUDA: PASS on both derivation packs, the 64 sampled words and the 96 vector lanes
against the Rust interpreter, and the 2^24 fingerprints equal to the Mac's, so the day program now agrees across
four compilers (Rust, Metal, Apple OpenCL, NVRTC). The build: 42 and 32 ms against x8's 40 and v2's 46 on the
loaded card (the unloaded x8 figure was 23 to 25 ms): the day program does not move the 5090's build beyond
noise, as on the Mac it moved it 7 ms. The hash rate: equal to v2 and x8 within the loaded noise (61 to 62
against 62), as it must be. THE NUMBER THAT MOVED: the NVRTC compile, 1,266 ms against 164 for x8 (+1.1 s per
pack), because kernel_bound.cu includes memhard.h and the 6,624-statement item function is compiled into every
hash kernel and every race variant: at 17 variants that is about +19 s of compile-ahead per epoch on the 5090
(approximate, 17 x 1.1 s; the race measured 232 to 300 ms for 17 variants today, epoch-length.md 6.1) unless the
item function is compiled once a day into its own module, which is the fix named in 5.5 and now a requirement
for the class, not an option: the 600-s epoch floor of the ladder leaves 38 s for the race today and this would
take half of it.
Consequences per tier: a 5090 owner pays about 1.1 s of compile once a day (the item library) plus, until the one-module fix lands, 1.1 s per epoch per variant; the Mac pays 0.75 s once a day; an AMD owner pays the OpenCL build (OWED, PC 1); the integrated tier's compile is unmeasured and its build was already the open problem.
5.5 The build per tier, with the integrated tier
| Card | Build at x8 | Build with the day program | Source |
|---|---|---|---|
| M5 Max, Metal | 22.1 ms | 29.0 ms (+32%) | 5.3, measure lock |
| RTX 5090, CUDA | 23 ms | section 5.4 | the PC 2 job |
| RX 9070 XT, OpenCL | 72 to 77 ms | OWED (PC 1) | |
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | about 55 to 94 s (approximate, mixer-x4.md 6.5: the iGPU's build is arithmetic-bound at x1 already, scaled x8 from 6.9 / 9.4 / 11.7 s) | about the same count of ops at a lower ILP: 55 to 120 s (approximate, unmeasured) | docs/plans/epoch-length.md 6.1 iGPU rows |
| gfx1036 beside WSL build jobs (PC 1) | about 7 to 17 min (approximate) | the same or worse (approximate) | epoch-length.md 6.1 |
| 8 GB-class discrete card (not owned, about a tenth of the 5090, approximate) | about 1 s | about 1 to 1.3 s (approximate) | scaled |
Consequences: nothing changes for a discrete card of any size on any vendor (the build is under a second), the Mac row measured; the integrated tier already misses the per-prepare rule at x8 and needs the per-day dataset reuse in the workers (0.3.12) or a restart per epoch, and the day program makes that need the same, not larger in kind; the per-day compile of the item function (0.75 s Metal; NVRTC on the 5090 in 5.4) lands once a day in the worker's day-cache build, not per epoch, unless the hash kernel's compile includes memhard.h, which on the CUDA worker it does (kernel_bound.cu includes it, and the variant race compiles 17 variants): the 5.4 job's nvrtc line is the number for that, and if it is large the fix is to compile the item function once per day into its own module.
6. PROPOSED spec text for 1.13.2: reserve entry R0, derive (the per-day item-derivation program)
Not written into docs/spec; it lives here until the founder's word. Named R0, ahead of R1 (mm8), because the reserve
is to be ordered by chip-unfriendliness (counter-asic-3.md item 6; mm8 last) and a derivation program is the most
chip-unfriendly entry the reserve can hold: it removes the fixed-function allowance of the recompute chip rather
than adding a family that chip can license.
Reserve entry R0,
derive(the per-day item-derivation program). Semantics: section 1.8.5 underderive_len = 736: the nine mixer slots of the item derivation (one before each of the 8 cache reads, one after the last) each run a straight-line program of 736 instructions drawn from the day key stream of 1.8.4 after its 40 draws, four draws per instruction (below(100)the form,below(15)the destination among the registers other than the chain,below(14)the third register among those other than the destination and the chain,next()the immediate:1 + low32 mod 31for the rotate forms,low32foraddcandxorc,low32 OR 1formulcandmulc2); the chain iss[0]at the start of each program and the previous destination after; the twelve forms and weights ofdocs/plans/counter-asic-3-derivation.mdsection 2.2, fixed; every form a bijection on the state; the 8 dependent cache reads, the mixer constants of the item init, the cache and the dataset mapping of 1.8.5 unchanged;mixer_multunused under R0. Acceptance test, per candidate, the next attempt on rejection (the stream continues): every register written in every round program; at least 8 distinct rotation amounts; per item at least 9,216 operations with constants folded, 10,368 as written and 1,152 multiplies (the x8 mixer's counts frommemhard::mixer: 72 x 128, 72 x 144, 72 x 16; the floors scale with the length). Edge vectors, each a hand-built item run on every vendor: item 0, item 1, item 2^28 - 1 and item 2^32 - 1 of the genesis day on a 2^16-word cache; a program whose first instruction is each of the twelve forms withd = 15,c = 0,b = 14,k = 31,i = 0xffffffff(odd for the multiplies) on the all-ones state and on the all-zero state (the wrap of every form); the day of the pinned packdr736-genesis(program fingerprint 463535d01511350d, dataset headvectors.json, 2^24 fingerprint 50e3eaa779da4f1e) and ofdr736-devnet-epoch0(771868df4e64d6ab, 9553f6d5c667205a). Unlock: at the start of era n = 2 (DAA 31,104,000), or earlier by the 90% signalling path of section 5.7, or at genesis if the verifier on a 2019-class core (O-1.14) reads under 10 ms per unit at 736, else at the length that does (368 measured at 2.69 ms on an M5 Max core); never by a release. The verifier procedure: the word-major interpreter ofigneum-pow/src/derive.rs(32 item states per batch, pair dispatch), 4.88 ms per unit on one M5 Max core (section 5.1), no JIT; a JIT is the named fallback (4.3). Vendor paths: none needed; every form is a single 32-bit integer statement on all three compilers (3.3).
7. The chip model row
Written into docs/analysis/chip-model-v3.md section 6 in the form of its section 2. Ops per hash: 128 items x
9,992 chip ops (the genesis day's draw; the floor 9,216) = 1,278,976 (floor 1,179,648); chip rate at 50 T op/s =
39.1 MH/s (floor 42.4); bare against 136.1 MH/s = 0.287x (floor 0.31x, the x8 row's figure, as the floor is x8's
count). The allowance rows: 1.0x 0.29x; 1.2x (ProgPoW's claim) 0.34x; 1.5x (cautious upper bound, approximate)
0.43x; 2x 0.57x; 3x (the fixed shape's, which no longer applies) 0.86x. Equal silicon (x 0.829): 0.24 / 0.29 /
0.36 / 0.48 / 0.71. The x8 row read 0.92x at 3x and 0.76x at equal silicon; at the same 0.31x bare the day program
takes the chip from 0.92x to 0.34x to 0.43x, which is the margin the item was for. dr368 (the fallback): 639,488
chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x, with less margin than x8 had at
3x and more than x4 had (1.84x).
8. What is unverified or owed
| Item | State |
|---|---|
RTX 5090 unloaded: the job's card-off did not take (the settings.json key carries the device index, nvidia:0:..., the api/cards key of the earlier jobs did not; the fix is to post both forms and confirm by the process list) so the 5090's absolute build and rate rows are loaded figures with the ratios valid (5.4a); the fingerprints and the compile are not affected |
a re-run after the key fix, on the next PC 2 slot |
| The item function as its own NVRTC module once a day (the +1.1 s per compile of 5.4a) | unimplemented; a requirement of the class before any activation |
| RX 9070 XT (PC 1) | OWED: PC 1 is the founder's desk today; the same job shape runs there with igneum-worker-opencl.exe --bench-pack when released |
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside M_r |
| The integrated tier's build with the day program | approximate (5.5); the gfx1036 measurement is a PC 2 OpenCL job, not run today (the one PC 2 job carries the 5090) |
| A 64-lane interpreter for pools (two units per batch) | unimplemented; it would cut the dispatch share for pool verifiers only |
| The NVRTC cost of memhard.h inside the per-epoch hash kernel compile and the variant race | the PC 2 job's nvrtc line; the fix, if large, is one module per day for the item function |
5.4b The RX 9070 XT rows (PC 1 job run-ca3-pc1-amd-derive-20261006, 6 October 2026, 17:23 to 17:29Z)
Beside the miners (ratios stand, absolutes are loaded-card figures). Both dr736 packs bit-exact on AMD OpenCL (gfx1201, AMD-APP 3683.0): fingerprints 9553f6d5c667205a (devnet epoch 0) and 50e3eaa779da4f1e (genesis), two passes each, self-tests PASS. OpenCL compile 2,070 and 2,092 ms per pack against 41 ms for mx8 (+2.0 s per pack; a once-a-day item module is required on AMD as on NVIDIA). Daily 1 GiB build 168 to 189 ms against 205 to 273 for mx8 in the same session. Rate ratio to mx8 0.982 and 1.006. The 9070 XT row of the owed list is closed; the chip model and the go / no-go are unchanged by it.