Counter ASIC 3.0 item 2: the design, the Mac measurements, the chip-model row and the PC 2 job

docs/plans/counter-asic-3-derivation.md (the design, the acceptance test, the interpreter, the allowance argument,
the measurements, the PROPOSED reserve entry R0 for 1.13.2, what is owed), docs/analysis/chip-model-v3.md section 6
(the per-day derivation rows at 1.0x to 3x allowances), the bench-log entry, relay/playbooks/ca3-derive-pc2.ps1
(one PC 2 job: self-fetched packs zip, the installed worker through NVRTC, the card off only under test with its
key from settings.json). Verifier 4.875 / 4.944 ms per unit on one M5 Max core under the measure lock against
x8's 2.061 / 2.063; Metal build 29 ms against 22; hash rate equal; bit-exact on Metal and Apple OpenCL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-06 08:50:01 +01:00
parent acb96ee8fb
commit dd5041b989
4 changed files with 657 additions and 0 deletions

View file

@ -98,3 +98,46 @@ order:
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point
under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had
no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.
## 6. The per-day derivation (item 2)
6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything
PROPOSED, a prototype behind load class `dr736`). The fixed-shape mixer of section 2's rows is replaced by nine
straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms,
the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item
derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program
of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a
32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's
constants in the wires. The counts are from the code (`memhard::mixer`: 144 ops per application as written, 128
with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1;
the day program's floor is those counts, so the bare row cannot fall below x8's.
| Row | Derivation | Chip ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) | Allowance 1.5x (cautious upper bound, approximate) | The old 3x (the fixed shape's; does not apply) | Equal silicon, SRAM deducted, at 1.2x / 1.5x |
|---|---|---|---|---|---|---|---|---|---|---|
| x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) | fixed mixer, 72 x 128 | 1,179,648 | 42.4 MH/s | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
| **dr736, the genesis day's draw** (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) | the day program, 9 x 736 instructions | 1,278,976 | 39.1 MH/s | 256 MiB | 128 / $46 | 0.29x | 0.34x | 0.43x | 0.86x | 0.29x / 0.36x |
| dr736 at the floor (a day whose draw sits exactly on the acceptance floor) | the day program | 1,179,648 | 42.4 | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
| dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) | the day program, 9 x 368 | 640,512 | 78.1 | 256 MiB | 128 / $46 | 0.57x | 0.69x | 0.86x | 1.72x | 0.57x / 0.71x |
| dr736 at year 4 (cache 512 MiB) | the day program | 1,278,976 | 39.1 | 512 MiB | 255 / $111 | 0.29x | 0.34x | 0.43x | 0.86x | 0.23x / 0.28x |
Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2
= 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357.
The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or
divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in;
with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function
is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector,
and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional
compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious
upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule
(every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no
intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave
items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument,
history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's
2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged;
the 5090's build and compile are the PC 2 job, the 9070 XT's OWED.
What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their
margins apply); no chip has been priced for its instruction store or its register files per item in flight; the
random ARX programs have had no cryptanalysis (the item 3 brief should name them beside `M_r`); the 2019-class
core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0
carries that verdict.

View file

@ -1961,3 +1961,84 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.
## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation
Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and
the chip row in `docs/plans/counter-asic-3-derivation.md` and `docs/analysis/chip-model-v3.md` section 6.
Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x
to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash
idea, `superscalar.cpp` read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op
count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build
and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average.
The construction (class `dr736`, `igneum-pow/src/derive.rs`): nine straight-line programs of 736 instructions per
item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64
stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by `c|1`, rotate-add, xor-rotate,
add-constant, xor-constant, the `M_r` form `(d ^ c) * odd`, `d * odd + c`, `d ^= c & b`, `d += c | b`), every
instruction reading the register the previous one wrote (the chain, `s[0]` first) and writing another, every form
a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations,
or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written,
1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624
instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major
interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT.
**Verifier per 32-lane unit, one core (`with-lock.sh measure`, one session 07:42:20 to 07:42:33 UTC, load
average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; `igneum-pow bench --seed igneum-genesis
--day 2026-10-03 --class <c> --warps 50`, two rounds, then the devnet seeds once):**
| Class | ms per unit, avg of 50 (round 1 / 2) | Worst cold of three | Against x8 | Ops per item (GPU / chip / mul) |
|---|---|---|---|---|
| v2 | 0.598 / 0.594 | 0.697 | | 1,296 / 1,152 / 144 |
| x8 (mx8, class v3) | 2.061 / 2.063 | 2.179 | 1 | 10,368 / 9,216 / 1,152 |
| dr736 | 4.875 / 4.944 | 5.241 | 2.37x | 10,659 / 9,992 / 1,461 |
| x8, devnet seeds | 2.078 | 2.155 | | |
| dr736, devnet seeds | 4.872 | 5.241 | 2.34x | 10,701 / 10,083 / 1,362 |
| dr368 (half length, the x4-equivalent fallback) | 2.692 | 2.898 | 1.31x | 5,350 / 5,004 / 752 |
The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core
figures. The interpreter's cost split (`examples/derive_perf.rs`, a functional run under the run lock, load 3.9 to
4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair
dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the
vector body (NEON, 1,180 `.4s` instructions in the binary).
**Bit-exactness (`with-lock.sh run`):** dr736-genesis on Metal (`packbench --batches 1 --batch-log2 24`) cache
FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint
2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (`igneum-bench-cl-dr736-genesis --bench-pack`) the
self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0
on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset
and on 2^24 outputs.
**Daily 1 GiB build and hash rate, Metal (`with-lock.sh measure`, the same session, `packbench --batches 2
--batch-log2 22 --group 256`, three rounds):**
| Pack | Compile (1 / 2 / 3) | Build, GPU ms (1 / 2 / 3) | MH/s GPU (1 / 2 / 3) |
|---|---|---|---|
| mx8-genesis (x8, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 |
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 |
**Chip model (chip-model-v3.md section 6):** 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at
50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no
longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a
cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.
**Consequences per tier.** The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms
(x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs
11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core
(2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The
build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below;
the 9070 XT is OWED (PC 1 is Josh's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier
already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to
94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the
same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache;
the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line
is the number to read. The hash rate: unchanged within 0.3% on the Mac, as the hash kernel only loads. Packs grow
by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.
**Go / no-go:** GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in
docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms;
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.

View file

@ -0,0 +1,415 @@
# Counter ASIC 3.0 item 2: a random item-derivation program per day
6 October 2026. Worker `derive` (branch `ca3-derive`), under the brief of `docs/plans/counter-asic-3.md` item 2 and
the history audit's addition 2 (`docs/analysis/asic-resistance-history.md` section 4.3). Everything here is
PROPOSED: a prototype behind a load class (`dr736`), measured on the Mac and on PC 2, written as a reserve entry for
spec 1.13.2 (section 6) that lives in this file until Josh's word. Nothing is published and no vector of class v2
or v3 moves (section 3.4).
What it does, in one line: the fixed-shape mixer `M_r` of spec 1.8.4, applied 72 times per item under class v3,
is replaced by nine straight-line programs of 736 instructions drawn once a day from the day key stream, so the chip
that holds the cache on die must run an arbitrary program instead of a wired pipeline; the 8 dependent cache reads
per item, the cache, the loads and the hash kernel are untouched.
## 0. The verdict first
| Gate | Number | Bar | Result |
|---|---|---|---|
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is Josh's desk today) |
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
core measurement (O-1.14) lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core
(pass) against about 12 ms on the approximate laptop row (fail). The class that passes both rows today is `dr368`
(2.69 ms), at the x4-equivalent op count, with the chip row at 0.58x bare. The way to the 736 figure under the gate
on a laptop is the JIT (section 4.3), which is out of scope and named with its risk.
## 1. Why: what the chip model says the fixed shape is worth
`docs/analysis/chip-model-v3.md` section 2 prices the on-die-cache recompute chip at 50 T op/s: class v3 (x8) costs
it 1,198,080 integer ops per hash (72 mixer applications x 128 items x about 130 ops), 41.7 MH/s, 0.31x the 5090's
136.1 MH/s bare, and 0.92x with "the 3x fixed-function factor (approximate, from memory: 2x to 5x is the usual
credit for a pipeline with no scheduling or divergence)". That credit is the mixer's fixed shape: the chip unrolls
the 72 applications into a wired pipeline with the day's constants baked in, no instruction fetch, no register
file, no operand muxes. RandomX's answer (`vendor/RandomX/src/superscalar.cpp` is not in this tree; read on
6 October 2026 from the upstream repository at commit 7607fb2 into the session scratchpad; the design argument is
`doc/design.md`, history section 2.4) is SuperscalarHash: the dataset item derivation is itself a random program
drawn from the cache key, 8 programs of about 450 instructions per item, so a light-mode chip "becomes a CPU".
RandomX's generator (`generateSuperscalar`, lines 653 to 850) schedules for a superscalar x86 core: it picks a
decode-buffer configuration per cycle, selects a source register that is ready at the cycle and a destination
that is not the source and was not last written by the same op group (`selectDestination`, line 495: no
"xor r,r2; xor r,r2", no "ror r,C1; ror r,C2", no two multiplies in a row on one register), and then computes the
program's ASIC latency as the longest dependency chain (lines 810 to 824) and sets the address register to the
register with the highest one. The item init is `rl[0] = (itemNumber + 1) * superscalarMul0`, the other seven
registers `rl[0] ^ superscalarAdd_i`, then per cache access: run the program, XOR the mix block in, next address
from the address register (`dataset.cpp` `initDatasetItem`, lines 164 to 190).
What carries over here and what does not: the per-day program, the acceptance by construction, the dependency
chain and the address register idea carry over; the x86 port scheduling does not (our verifier interprets the
program for 32 items at once and our miners compile it for a GPU, so the schedule that matters is the GPU's), and
the latency bound RandomX relies on (the program's critical path against DRAM) is replaced by ours: the 8
dependent cache reads per item, which are untouched.
## 2. The design
### 2.1 The draw
One SplitMix64 stream seeded with `K[0] | (K[1] << 32)` (the day key, spec 1.8.1), the 40 draws of 1.8.4
(`ROT`, `MUL`, `RC`) first, exactly as today (the item init `s[8 + i] = t * MUL[i] + RC[i]` still uses them), then
the program: `DERIVE_PROGRAMS = 9` round programs (one before each of the 8 cache reads, one after the last) of
`derive_len` instructions each, four draws per instruction in a fixed order:
| Draw | Range | Sets |
|---|---|---|
| `below(100)` | op roll | the form, by the weight table of 2.2 (cumulative) |
| `below(15)` | destination roll | `d`: the 15 registers other than the chain `c`, in ascending order (roll >= c adds one) |
| `below(14)` | third-register roll | `b`: the 14 registers other than `d` and `c`, ascending (two skips); used by `andx` and `orx`, consumed by every form |
| `next()` | the immediate | `k = 1 + (low32 mod 31)` for the rotate forms; `imm = low32` for `addc`, `xorc`; `imm = low32 OR 1` for `mulc`, `mulc2`; consumed by every form |
The chain register `c` is `s[0]` (the address word) at the start of each round program and the destination of the
previous instruction after it. Every instruction reads `c` and writes `d != c`; the next instruction's chain is
`d`. So no two instructions of a program can run in parallel (the "every instruction consumes the newest result"
rule of the brief, SuperscalarHash's chain made strict), and no two consecutive instructions write one register,
which is what removes the mergeable pairs SuperscalarHash's `selectDestination` guards against. Four draws per
instruction whether the form uses them or not, so the stream position of every draw is fixed by its index and a
future change to one form's draw changes no other draw (the rule 1.13.1 follows for `epoch_len`). A rejected
candidate (2.4) is followed by the next: the stream continues, `attempt + 1`, as the program generator of 1.4.6.
Source: `igneum-pow/src/derive.rs` (`DeriveProgram::draw_candidate`, `draw`), `memhard.rs` (`MixParams::with_shape`).
### 2.2 The instruction set: twelve two-register forms, fixed at genesis
`c` the chain, `d` the destination, `b` the third register, `k` in 1..31, `i` a 32-bit constant (odd for the
multiplies). All arithmetic modulo 2^32, no division, no float, no data-dependent branch. Every form is a bijection
on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply, or a rotation of itself; `c`
and `b` are not written), so a program loses no entropy, the property `M_r` has.
| Form | Semantics | Weight (percent) | GPU ops | Chip ops | Multiply |
|---|---|---|---|---|---|
| `add` | `d += c` | 14 | 1 | 1 | |
| `sub` | `d -= c` | 10 | 1 | 1 | |
| `xor` | `d ^= c` | 14 | 1 | 1 | |
| `mul` | `d *= (c OR 1)` | 10 | 2 | 1 (the OR is a wire) | yes |
| `rot` | `d = rotl(d, k) + c` | 10 | 2 | 2 | |
| `xrot` | `d = rotl(d ^ c, k)` | 10 | 2 | 2 | |
| `addc` | `d += c + i` | 6 | 2 | 2 | |
| `xorc` | `d ^= c ^ i` | 6 | 2 | 2 | |
| `mulc` | `d = (d ^ c) * i` (the per-word form of `M_r`, the chain in place of the round constant) | 8 | 2 | 2 | yes |
| `mulc2` | `d = d * i + c` | 4 | 2 | 2 | yes |
| `andx` | `d ^= (c AND b)` | 4 | 2 | 2 | |
| `orx` | `d += (c OR b)` | 4 | 2 | 2 | |
Mean 1.62 GPU ops and 1.52 chip ops per instruction, 22% multiplies. `and` and `or` enter only as `andx` and
`orx` (a destructive `d &= c` would lose bits; the XOR and add of a conjunction keep `d` invertible). The forms are
the brief's set (add, sub, mul-lo, xor, rotate by 1..31, and, or, the ARX-multiply forms of `M_r`); `mulhi` is in
the lottery hash's families and bit-exact on the three vendors, and is left out of the derivation on purpose so
every form is one that the three compilers lower to a single integer instruction (section 3.3).
### 2.3 Length and the floors: the x8-equivalent op count
The coordinator's rule for item 2: the total op count per item equals or exceeds today's x8 count, so the chip's
budget row does not fall. The x8 mixer counted from the code (`memhard::mixer`): 16 x (xor, add, mul) + 8 quarter
rounds x 12 = 144 ops as written, 128 with the `RC[i] + rk` adds hoisted as constants (the chip and every compiler
do that), 16 multiplies; 72 applications per item = 10,368 ops as written, 9,216 hoisted, 1,152 multiplies
(`chip-model-v3.md` section 1 prices 130 per application from the spec text; the item 3 worker counted the same
144 / 128, coordinator's note of 6 October). The floors are those three. `derive_len = 736` gives 6,624
instructions per item, expected 10,731 GPU ops, 10,068 chip ops and 1,457 multiplies (the genesis day draws
10,659 / 9,992 / 1,461; the devnet day 10,701 / 10,083 / 1,362), 7 to 22 standard deviations above the floors, so a
rejection on a floor is a rare event and the test exists for the degenerate class. The class name is `dr736`; the
half-length `dr368` (the x4-equivalent) is measured beside it as the fallback with its floors scaled.
### 2.4 The acceptance test
`DeriveProgram::check`, on every candidate; a rejection draws the next attempt from the stream:
| Test | Rejects | Expected rate at 736 |
|---|---|---|
| every register written in every round program | a register no instruction of a round program writes (its init word would never enter the chain within that round) | 16 x 9 x (14/15)^736 = under 10^-20 |
| at least 8 distinct rotation amounts across the item's programs | all rotations equal (the brief's degenerate draw; the mixer's "ROT draw of eight equal values is possible and untested", spec 1.8.4) | about 1,300 rotate forms over 31 values: never |
| chip ops >= 9,216, GPU ops >= 10,368, multiplies >= 1,152 per item (scaled to the length) | an op draw under the x8 count | 22, 9 and 9 standard deviations below the mean: never |
| structural (asserted, hold by construction): `d != c`, `b` distinct from both, `k` in 1..31, odd multiplier constants | a generator bug | n/a |
The generator panics after 64 rejected candidates in a row (`MAX_ATTEMPTS`), which the rates above put beyond
any day the chain will see; the panic is the right failure (a node that cannot derive the day's program cannot
verify, and must say so rather than guess). The weights are fixed at genesis; only the order, the registers and
the constants are drawn, so the family mix of a program cannot be steered by the draw.
### 2.5 The item, with the program in place
Spec 1.8.5 under the derivation class (`derive_len` nonzero, `mixer_mult` unused):
```
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
s = P_r(s) round program r, the chain starting at s[0]
a = s[0] AND (2^(C - 4) - 1) cache line index, as today
s[i] = s[i] XOR cache[line a][i] for i in 0..15
s = P_8(s)
item(t) = s
```
The 8 dependent cache reads per item are exactly today's: the address of read `r` is `s[0]` after program `r`, and
`s[0]` depends on every earlier read through the chain (every round program writes every register, 2.4, and the
XOR of the line into all 16 words feeds the next program). The verifier's latency part (8 dependent misses per
item, overlapped across the up to 32 items of a load, spec 1.11) is unchanged, which the measurement shows: the
difference against x8 is the ALU part only (section 5.1).
## 3. The prototype
### 3.1 Where it lives
| Item | Where |
|---|---|
| `DOp`, `DInstr`, `DeriveProgram` (draw, check, counts, fingerprint), the SoA interpreter `run_round` (pair dispatch), the scalar reference `run_round_scalar`, the text forms `instr_text` and `instr_line` | `igneum-pow/src/derive.rs` |
| `Shape::derive_len`, `Shape::is_derived`, `MixParams::derive` (drawn after the 40 mixer draws), `derive_items` dispatching to `derive_items_program` | `igneum-pow/src/memhard.rs` |
| `LoadClass::derive_len`, `LoadClass::DR736`, `with_derive`, parse and name `dr<len>`, the program id (`derive/` + the length) | `igneum-pow/src/generator.rs` |
| `mh_round_0..8` and the program-driven `mh_item` in memhard.h, memhard.metal and kernel.cl; `IGNEUM_DERIVE_*` in program.h; `derive_len`, the op mix, the floors' counts and the nine programs (one line per instruction) in program.json | `igneum-pow/src/emit.rs` |
| `--class dr736` (or any `dr<len>`) on every command; the bench prints the program's counts | `igneum-pow/src/main.rs` |
| `examples/derive_perf.rs`: the interpreter's cost per instruction per batch, drawn program against uniform programs | `igneum-pow/examples/` |
| Packs `dr736-genesis` (seed igneum-genesis, day 2026-10-03, program id 72c1d8048aef9542, program fingerprint 463535d01511350d) and `dr736-devnet-epoch0` (the devnet epoch 0 and day seeds, 7f4a5ca0a3637820, 771868df4e64d6ab); generator 2 with the class in the id, the x4-record shape, not a class v3 pack | `proto-cuda/packs-ca3-derive/` |
| Tests: `src/derive.rs` (5), `tests/derive.rs` (7: by hand on a small cache, batches, v2 and v3 untouched, the stream and the class, determinism and the pack text, stats beside x8, the text forms against the scalar reference, the word path) | `igneum-pow` |
| The PC 2 job | `relay/playbooks/ca3-derive-pc2.ps1` (section 5.4) |
### 3.2 The verifier without a JIT: the word-major interpreter
The verifier derives up to 32 distinct items per load (one per lane of the unit, `MemhardCpu::fetch`). The
interpreter keeps the 32 item states word-major (`st[reg][lane]`, 2 KiB) and runs each instruction across the
whole batch in one straight loop the compiler vectorises (NEON `add.4s`, `mul.4s`, `ushl.4s` and so on: 1,180
such instructions in the example binary), so the dispatch is paid once per instruction per batch, not per item;
the cache reads of the batch are issued together after each round program, as the fixed-mixer loop does, so the 8
dependent misses of independent items overlap. The dispatch is on PAIRS of instructions (144 arms, one indirect
branch per two instructions): a drawn op sequence is random, the predictor misses most dispatches, and pairing
halves the misses per instruction. Measured with `examples/derive_perf.rs` (a functional run, load average 3.9 to
4.9): 10.7 ns per instruction per batch cold, 7.18 warm with single dispatch, 4.98 with pair dispatch; a uniform
program of one form (predictable dispatch) 3.4 to 4.3 ns, so the body is about 3.5 ns and the remaining dispatch
cost about 1.5 ns. 6,624 x 4.98 ns x 128 batches = 4.2 ms per unit of interpreter time; the measured 4.88 ms
includes the latency part and the transposes.
### 3.3 Bit-exactness on three vendors, by construction
Each form is one C statement on `uint` with `+`, `-`, `^`, `*`, `|`, `&` and the memhard core's `mh_rotl` (a
shift pair, `n` in 1..31 at every call site), the same text in Metal, CUDA C and OpenCL C, every operand a 32-bit
unsigned integer: the same argument as spec 1.14 for the lottery hash's families, which have run bit-exact on the
three vendors since 4 October. The program is emitted as nine functions of straight-line statements (196 KB of
memhard.h per day); NVRTC, the Metal compiler and the OpenCL compilers see no loop, no branch and no call inside a
round program. Measured: section 5.2 (Metal and Apple OpenCL), 5.4 (CUDA).
### 3.4 The v2 and v3 paths are untouched
`Shape::derive_len` is 0 and `LoadClass::derive_len` is 0 on `V2`, `MX4`, `MX8` and `V3_CLASS`; `MixParams::derive`
is `None`; `derive_items` takes the fixed-mixer loop as before; the emitter's text for a shape without a program is
unchanged. `cargo test -p igneum-pow` (commit acb96ee): 58 lib, 7 derive, 4 mixer, 19 packs (every pinned pack of
v2 and v3 regenerated and compared byte for byte), 7 scratch, all green.
## 4. What the chip keeps, and the fallbacks
### 4.1 The fixed-function allowance after the change
The chip of `chip-model-v3.md` now has to execute, per item, 6,624 instructions from a 12-form set over a
16-entry register file, with three operand fields and a constant per instruction, in an order and with operands
that change every day. That is a sequencer: an instruction store (6,624 x 9 bytes = 60 KB per day program, in
SRAM beside the cache mirror), a register file with three read ports and one write port, a 32-bit ALU with a
multiplier and a barrel rotator, operand muxes, and a program counter. A GPU streaming multiprocessor is the same
machine with a wider register file and a warp scheduler. What the chip keeps over the GPU: no warp scheduler, no
operand collector, no instruction cache hierarchy for a 1 MB kernel (the GPU's day program is about 60,000 SASS
instructions per thread; whether the 5090's instruction cache holds it is in the build time of 5.4), no graphics
or float units idle on the die, and the day's constants folded into the instruction store. What it loses: the
wired pipeline (the mixer's 72 applications as 72 stages with the constants in the wires, no fetch, no register
file, no crossbar), which is the thing the 3x credit paid for. The chain rule adds a second loss: with no
intra-item parallelism a single engine finishes one instruction per cycle at best and its multiplies serialise
unless it interleaves items, which costs a register file per item in flight (RandomX's light-mode argument,
history 2.4: a chip paying "760 cycles and 1,240 multiplies per item"). ProgPoW claimed 1.1x to 1.2x for exactly
this kind of chip ("conventional compute chips gain little on ProgPoW", Bob Rao's hardware audit, history 2.4,
[S70]; EIP-1057's own claim 1.1x to 1.2x, [S67]). So the allowance this document carries is 1.2x (ProgPoW's
claimed range, cited) with 1.5x as the cautious upper bound (approximate, mine), against the 3x of the fixed shape
(approximate, from memory, M16). Section 7 prices all of 1.0x, 1.2x, 1.5x, 2x and 3x.
### 4.2 What it does not change
The partial-store chip (item 1): a chip that stores the dataset and never derives items pays nothing for the
program; item 1's rows stand on their own. The cryptanalysis question (item 3) changes shape: instead of one
fixed `M_r` to attack, the attacker gets a fresh random ARX program every day, which is RandomX's bet and is
untested here; the day's programs can be audited by the same tools as random ARX ciphers, and the acceptance test
is the place to add a structural rule if one is found. The era draws and the cache growth are untouched.
### 4.3 The JIT, named as the fallback with its risk
A per-day JIT for the verifier (emit NEON or AVX2 code for the 32-lane batch, the chain row kept in registers
across instructions since `c` is always the row just written) would remove the 1.5 ns dispatch and about a third
of the 3.5 ns body (the chain row's loads), about 2.5 ms per unit on this core, the x8 figure. Its risk: a code
generator in the consensus path on two architectures (arm64, x86-64) whose output must equal the interpreter's
bit for bit, writable-executable memory in a node and in every pool verifier, a new attack surface the history's
lesson 9 says to audit before launch, and a second implementation per platform to keep in lockstep. Out of scope
for this item; it is the route to 736 under the gate on a laptop core if the measurement of O-1.14 confirms the
approximate row, and `dr368` is the route that needs no JIT.
## 5. Measurements
All under the locks of the brief; every row says its lock and the load average. Machine: Apple M5 Max, 64 GiB,
Darwin 25.6.0. Commit acb96ee (the code and the packs), packbench built from this worktree.
### 5.1 The verifier per unit, one M5 Max core (`with-lock.sh measure`, one session, 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end)
Script `measure-ca3-derive.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03
--class <v2|mx8|dr736> --warps 50`, two rounds, then the devnet seeds (`--epoch-hex edc4fa84...fb07 --day-hex
69676e65756d2d6461792ffa50000000000000`) and `dr368` once.
| Class | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | Against x8 | Items per unit | Ops per item (GPU / chip / multiplies) |
|---|---|---|---|---|---|---|
| v2 (igneum-genesis) | 0.598 / 0.594 | 0.697 | 1 | | 4,096 | 1,296 / 1,152 / 144 (9 x 144, hoisted 9 x 128) |
| x8, mx8 (class v3) | 2.061 / 2.063 | 2.179 | 3.46x | 1 | 4,096 | 10,368 / 9,216 / 1,152 |
| **dr736** | **4.875 / 4.944** | **5.241** | 8.2x | 2.37x | 4,096 | 10,659 / 9,992 / 1,461 |
| x8, the devnet seeds | 2.078 | 2.155 | | | 4,095 to 4,096 | |
| dr736, the devnet seeds | 4.872 | 5.241 | | 2.34x | 4,096 | 10,701 / 10,083 / 1,362 |
| dr368 (the x4-equivalent fallback) | 2.692 | 2.898 | 4.5x | 1.31x | 4,096 | 5,350 / 5,004 / 752 |
Reading: the v2 row reads 0.59 to 0.60, the quiet readwidth night's 0.604 to 0.626 and 6.4a's 0.607 to 0.611, so
this session is a quiet-core figure and no scaling applies. At the same chip-op count as x8 (plus 8%), the
interpreter costs 2.4x what the compiled mixer costs: 4.98 ns per instruction per batch (3.2), of which about
1.5 ns is dispatch and 3.5 ns the vector body, against the mixer's compiled straight-line loop over the same
batch. The latency part is the same in every row (the 8 dependent misses per item; the dr368 row at half the
instructions saves 2.2 ms of the 4.9, which puts the latency-and-transpose share at about 0.5 ms). The 10 ms gate
keeps 5.1 ms (worst cold 4.76 ms) at 736 on this core and 7.3 ms at 368.
### 5.1a Verification throughput per tier (the form of mixer-x4.md 6.5)
| Figure | x8 (this session) | dr736 | dr368 | Note |
|---|---|---|---|---|
| ms per unit, quiet M5 Max core (measured, this session) | 2.06 | 4.88 | 2.69 | |
| ms per unit, 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) | 5.2 | 12.2 | 6.7 | the figure that fixes the gate is a measurement, not this row |
| Shares per second per core (quiet M5 Max) | 485 | 205 | 372 | |
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second) | 4.5 | 10.7 | 5.9 | a pool verifying two units at once would halve the dispatch share (3.2); a 64-lane interpreter is a follow-up, unmeasured |
| Node: worst cold single unit (per block) | 2.2 ms | 5.2 ms | 2.9 ms | a block's verification stays under the 1 s block time by 190x |
| IBD over 108,000 headers on one core | 3.7 min | 8.8 min | 4.8 min | laptop (approximate): 9.4 / 22 / 12 min |
| Margin left under the 10 ms gate (worst cold, this core) | 7.8 ms | 4.8 ms | 7.1 ms | on the laptop row (approximate): 4.8 / none (over by 2.2) / 3.3 ms |
Consequences per tier: every miner tier is untouched by the verifier (the miner never runs it); a pool operator
pays 2.4x the cores at 736 (11 cores for a 22,000-member pool against 4.5) or 1.3x at 368; a node on any 2026 core
verifies a block in 5 ms; a node on a 2019-class laptop core is the open question (O-1.14), and at 736 the
approximate row says it misses the gate, so 736 does not go genesis-live on an approximation, and 368 passes it.
### 5.2 Bit-exactness on the Mac (`with-lock.sh run`, 08:38 to 08:41 local)
`packbench --pack <dir> --batches 1 --batch-log2 24 --group 256` (Metal, built from this worktree) and
`igneum-bench-cl-dr736-genesis --bench-pack --pack <dir> --batches 1 --batch-log2 24` (Apple OpenCL, `proto-opencl/
build.sh` on the pack through a temporary link). Vectors are the Rust interpreter's.
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word [MASK], 64 samples | Vectors | Fingerprint 2^24 | Compile | 1 GiB build, GPU ms (run lock, indicative) |
|---|---|---|---|---|---|---|---|
| dr736-genesis | Metal | 48c4f5bf24166b2e PASS | head and last PASS (packbench checks no samples) | 3/3 standalone, 3/3 in batch | 50e3eaa779da4f1e | 784 ms (276 on the second process, 1 ms once the shader cache has it) | 38.2 / 38.3 |
| dr736-genesis | Apple OpenCL | PASS (head, last line, FNV) | head, word [268435455], 64 samples PASS | 96 of 96 lanes | 50e3eaa779da4f1e | (in the 385 ms prepare) | 57 ms wall |
| dr736-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 9553f6d5c667205a | 763 ms | 37.4 |
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the derived dataset (head, word [MASK], the 64
samples through OpenCL), on every vector lane and on the 2^24-output fingerprint of dr736-genesis across both
compilers; the devnet-seed pack agrees on Metal. The Apple OpenCL compile of a 6,624-statement item function
is inside a 385 ms prepare; the Metal compile is 0.75 to 0.8 s cold per day (the memhard library is the day's,
compiled once a day, not per epoch) and 1 ms from the shader cache.
### 5.3 The daily build and the hash rate on the M5 Max (`with-lock.sh measure`, the session of 5.1, three rounds)
`packbench --pack <dir> --batches 2 --batch-log2 22 --group 256`, mx8-genesis then dr736-genesis, three rounds.
| Pack | Compile (round 1 / 2 / 3) | 1 GiB build, GPU ms (round 1 / 2 / 3) | MH/s GPU (round 1 / 2 / 3) | Vectors, self-tests |
|---|---|---|---|---|
| mx8-genesis (class v3, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | 3/3 + 3/3, PASS |
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | 3/3 + 3/3, PASS |
Reading: the build is 29 ms against 22 (+32%, +7 ms): the Mac's build was latency-bound at x1, x4 and x8 (21 to
22 ms at every multiplier, mixer-x4.md 6.4) and the serial chain per thread now shows (no intra-item ILP for the
compiler to schedule), still 34x under the 1 s bar. The hash rate is the x8 rate within 0.3% (27.06 to 27.13
against 27.08 to 27.16), as it must be: the hash kernel only loads. Consequences per tier: a daily build of 29 ms
on Apple silicon costs nothing to any tier; the 5090's figure is section 5.4, the 9070 XT's is OWED (PC 1 not
released today; its x8 build was 72 to 77 ms and arithmetic-bound at none of x1, x4, x8, so the chain's cost there
is the open number); the integrated tier is the one to watch (5.5).
### 5.4 RTX 5090, PC 2 (one job, `relay/playbooks/ca3-derive-pc2.ps1`)
PENDING at the time of writing: the job waits for `/tmp/igneum-devnet/pc2-ca3.clear` (the proving agent's
30-minute measurement) and the `pc2-ca3.lock`. The job downloads the packs zip itself (sha256
aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676: dr736-genesis, dr736-devnet-epoch0, mx8-genesis,
v2-genesis-mh), runs the installed app's igneum-worker-cuda.exe (NVRTC compiles each pack's own text) with the
NVIDIA card off in the app only under test (its key from settings.json), `--check` for the `nvrtc .. cache ..
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
The rows land in the bench-log entry when the closing report is read.
### 5.5 The build per tier, with the integrated tier
| Card | Build at x8 | Build with the day program | Source |
|---|---|---|---|
| M5 Max, Metal | 22.1 ms | 29.0 ms (+32%) | 5.3, measure lock |
| RTX 5090, CUDA | 23 ms | section 5.4 | the PC 2 job |
| RX 9070 XT, OpenCL | 72 to 77 ms | OWED (PC 1) | |
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | about 55 to 94 s (approximate, mixer-x4.md 6.5: the iGPU's build is arithmetic-bound at x1 already, scaled x8 from 6.9 / 9.4 / 11.7 s) | about the same count of ops at a lower ILP: 55 to 120 s (approximate, unmeasured) | `docs/plans/epoch-length.md` 6.1 iGPU rows |
| gfx1036 beside WSL build jobs (PC 1) | about 7 to 17 min (approximate) | the same or worse (approximate) | epoch-length.md 6.1 |
| 8 GB-class discrete card (not owned, about a tenth of the 5090, approximate) | about 1 s | about 1 to 1.3 s (approximate) | scaled |
Consequences: nothing changes for a discrete card of any size on any vendor (the build is under a second), the
Mac row measured; the integrated tier already misses the per-prepare rule at x8 and needs the per-day dataset
reuse in the workers (0.3.12) or a restart per epoch, and the day program makes that need the same, not larger in
kind; the per-day compile of the item function (0.75 s Metal; NVRTC on the 5090 in 5.4) lands once a day in the
worker's day-cache build, not per epoch, unless the hash kernel's compile includes memhard.h, which on the CUDA
worker it does (kernel_bound.cu includes it, and the variant race compiles 17 variants): the 5.4 job's nvrtc line
is the number for that, and if it is large the fix is to compile the item function once per day into its own
module.
## 6. PROPOSED spec text for 1.13.2: reserve entry R0, `derive` (the per-day item-derivation program)
Not written into `docs/spec`; it lives here until Josh's word. Named R0, ahead of R1 (mm8), because the reserve
is to be ordered by chip-unfriendliness (counter-asic-3.md item 6; mm8 last) and a derivation program is the most
chip-unfriendly entry the reserve can hold: it removes the fixed-function allowance of the recompute chip rather
than adding a family that chip can license.
> Reserve entry R0, `derive` (the per-day item-derivation program). Semantics: section 1.8.5 under
> `derive_len = 736`: the nine mixer slots of the item derivation (one before each of the 8 cache reads, one after
> the last) each run a straight-line program of 736 instructions drawn from the day key stream of 1.8.4 after its
> 40 draws, four draws per instruction (`below(100)` the form, `below(15)` the destination among the registers
> other than the chain, `below(14)` the third register among those other than the destination and the chain,
> `next()` the immediate: `1 + low32 mod 31` for the rotate forms, `low32` for `addc` and `xorc`, `low32 OR 1` for
> `mulc` and `mulc2`); the chain is `s[0]` at the start of each program and the previous destination after; the
> twelve forms and weights of `docs/plans/counter-asic-3-derivation.md` section 2.2, fixed; every form a
> bijection on the state; the 8 dependent cache reads, the mixer constants of the item init, the cache and the
> dataset mapping of 1.8.5 unchanged; `mixer_mult` unused under R0. Acceptance test, per candidate, the next
> attempt on rejection (the stream continues): every register written in every round program; at least 8 distinct
> rotation amounts; per item at least 9,216 operations with constants folded, 10,368 as written and 1,152
> multiplies (the x8 mixer's counts from `memhard::mixer`: 72 x 128, 72 x 144, 72 x 16; the floors scale with the
> length). Edge vectors, each a hand-built item run on every vendor: item 0, item 1, item 2^28 - 1 and item 2^32 - 1
> of the genesis day on a 2^16-word cache; a program whose first instruction is each of the twelve forms with
> `d = 15`, `c = 0`, `b = 14`, `k = 31`, `i = 0xffffffff` (odd for the multiplies) on the all-ones state and on the
> all-zero state (the wrap of every form); the day of the pinned pack `dr736-genesis` (program fingerprint
> 463535d01511350d, dataset head `vectors.json`, 2^24 fingerprint 50e3eaa779da4f1e) and of `dr736-devnet-epoch0`
> (771868df4e64d6ab, 9553f6d5c667205a). Unlock: at the start of era n = 2 (DAA 31,104,000), or earlier by the 90%
> signalling path of section 5.7, or at genesis if the verifier on a 2019-class core (O-1.14) reads under 10 ms per
> unit at 736, else at the length that does (368 measured at 2.69 ms on an M5 Max core); never by a release. The
> verifier procedure: the word-major interpreter of `igneum-pow/src/derive.rs` (32 item states per batch, pair
> dispatch), 4.88 ms per unit on one M5 Max core (section 5.1), no JIT; a JIT is the named fallback (4.3). Vendor
> paths: none needed; every form is a single 32-bit integer statement on all three compilers (3.3).
## 7. The chip model row
Written into `docs/analysis/chip-model-v3.md` section 6 in the form of its section 2. Ops per hash: 128 items x
9,992 chip ops (the genesis day's draw; the floor 9,216) = 1,278,976 (floor 1,179,648); chip rate at 50 T op/s =
39.1 MH/s (floor 42.4); bare against 136.1 MH/s = 0.287x (floor 0.31x, the x8 row's figure, as the floor is x8's
count). The allowance rows: 1.0x 0.29x; 1.2x (ProgPoW's claim) 0.34x; 1.5x (cautious upper bound, approximate)
0.43x; 2x 0.57x; 3x (the fixed shape's, which no longer applies) 0.86x. Equal silicon (x 0.829): 0.24 / 0.29 /
0.36 / 0.48 / 0.71. The x8 row read 0.92x at 3x and 0.76x at equal silicon; at the same 0.31x bare the day program
takes the chip from 0.92x to 0.34x to 0.43x, which is the margin the item was for. dr368 (the fallback): 639,488
chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x, with less margin than x8 had at
3x and more than x4 had (1.84x).
## 8. What is unverified or owed
| Item | State |
|---|---|
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
| RX 9070 XT (PC 1) | OWED: PC 1 is Josh's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |
| The integrated tier's build with the day program | approximate (5.5); the gfx1036 measurement is a PC 2 OpenCL job, not run today (the one PC 2 job carries the 5090) |
| A 64-lane interpreter for pools (two units per batch) | unimplemented; it would cut the dispatch share for pool verifiers only |
| The NVRTC cost of memhard.h inside the per-epoch hash kernel compile and the variant race | the PC 2 job's nvrtc line; the fix, if large, is one module per day for the item function |

View file

@ -0,0 +1,118 @@
# Igneum run job: Counter ASIC 3.0 item 2, the per-day item-derivation program (docs/plans/counter-asic-3-derivation.md),
# on PC 2's RTX 5090 (machine 1ccfe586), 6 October 2026. ONE job carries everything (the PC 2 rule of the ca3 brief):
# it downloads the packs zip itself from the downloads host (sha256 checked), finds the installed app's
# igneum-worker-cuda.exe (NVRTC compiles each pack's own kernel text, so no new worker build is needed), switches the
# NVIDIA card off in the app ONLY while the packs run (the card key from the app's settings.json, never /api/state;
# restored after with the settings it had), and never quits, restarts or updates the installed app. The numbers this
# job is for: per pack, the worker's own `nvrtc .. cache .. dataset .. ms` line (the daily 1 GiB build on the 5090),
# the self-test (bit-exactness: cache FNV, dataset head, word [MASK], 64 samples, 96 vector lanes against the Rust CPU
# interpreter), the 2^24 fingerprint at base nonce 0 (against the Mac: dr736-genesis 50e3eaa779da4f1e,
# dr736-devnet-epoch0 9553f6d5c667205a, mx8-genesis 7c28cfb06c5c65a9, v2-genesis-mh 25f96e7dce90bd4e) and the hash rate.
# Packs: dr736-genesis, dr736-devnet-epoch0 (the derivation class), mx8-genesis (the x8 control), v2-genesis-mh (the v2
# control). Every result line starts with RESULT. The placeholder __DL_BASE__ is substituted at publish time; the
# downloads token never enters the repository.
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
# everything lives in this job's own folder (the wiped-jobs-folder rule of tools/ci/kit-path-check.sh: nothing of
# another job's is reached)
$work = Join-Path $env:IGNEUM_JOB_DIR 'ca3-derive'
New-Item -ItemType Directory -Force -Path $work | Out-Null
$zipUrl = '__DL_BASE__/igneum-ca3-derive-packs.zip'
$zipSha = 'aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676'
$zip = Join-Path $work 'igneum-ca3-derive-packs.zip'
Write-Output "RESULT job ca3-derive-pc2 start $(Stamp) machine $env:COMPUTERNAME"
# ---- the packs: downloaded by this job, sha256 checked, expanded fresh ----
try {
[Net.ServicePointManager]::SecurityProtocol = [Net.SecurityProtocolType]::Tls12
Invoke-WebRequest -Uri $zipUrl -OutFile $zip -TimeoutSec 120 -Headers @{ 'Cache-Control' = 'no-cache' }
} catch { Write-Output ("RESULT error download: " + $_.Exception.Message); exit 2 }
$got = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower()
if ($got -ne $zipSha) { Write-Output "RESULT error zip sha256 $got expected $zipSha"; exit 2 }
Write-Output "RESULT zip sha256 $got size $((Get-Item $zip).Length) OK"
$packs = Join-Path $work 'packs-ca3-derive'
if (Test-Path $packs) { Remove-Item -Recurse -Force $packs }
Expand-Archive -Path $zip -DestinationPath $work -Force
$packList = @('v2-genesis-mh', 'mx8-genesis', 'dr736-genesis', 'dr736-devnet-epoch0')
foreach ($pk in $packList) {
$d = Join-Path $packs $pk
if (-not (Test-Path (Join-Path $d 'memhard.h'))) { Write-Output "RESULT error pack $pk missing after extract"; exit 2 }
Write-Output ("RESULT pack $pk memhard.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'memhard.h')).Hash.ToLower() + " program.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'program.h')).Hash.ToLower())
}
# ---- the worker: the installed app's igneum-worker-cuda.exe, run in place (its NVRTC DLLs sit beside it) ----
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe'; exit 2 }
$cuda = Join-Path $inst 'igneum-worker-cuda.exe'
Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) nvrtc_dlls $((Get-ChildItem $inst -Filter 'nvrtc*.dll').Count)"
$help = (& $cuda --help 2>&1 | Out-String)
$hasBench = $help -match '--bench'
Write-Output "RESULT worker-cuda has --bench: $hasBench"
if (-not $hasBench) { Write-Output 'RESULT note the installed worker has no --bench; --check gives the build time and the self-test, no fingerprint and no hash rate' }
# ---- the app: the NVIDIA card's key and settings from settings.json; switched off through POST api/cards only ----
$appDir = $env:IGNEUM_APP_DIR
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
$urlFile = Join-Path $appDir 'app.url'
$url = $null
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
$sj = Join-Path $appDir 'settings.json'
$cardKey = $null; $cardPref = $null
if (Test-Path $sj) {
try {
$settings = Get-Content -LiteralPath $sj -Raw | ConvertFrom-Json
if ($settings.cards) {
foreach ($p in $settings.cards.PSObject.Properties) { if ($p.Name -like 'nvidia:*') { $cardKey = $p.Name; $cardPref = $p.Value; break } }
}
} catch { Say ("settings.json: " + $_.Exception.Message) }
}
if ($cardKey) {
Write-Output ("RESULT card " + $cardKey + " enabled=" + $cardPref.enabled + " identities=" + $cardPref.identities + " power_pct=" + $cardPref.power_pct + " (settings.json)")
} else {
Write-Output 'RESULT card none in settings.json (no nvidia:* entry); the app keeps mining on the card and the numbers carry that load'
}
$cardOff = $false
if ($cardKey -and $url) {
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
$body = @{ cards = @(@{ key = $cardKey; enabled = $false; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; $cardOff = $true; Say 'card off requested' } catch { Write-Output ("RESULT error api/cards off: " + $_.Exception.Message) }
# wait for the worker process to go, up to 90 s (the process list, not /api/state)
$t = 0
while ($t -lt 90) {
Start-Sleep -Seconds 5; $t += 5
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
if (-not $w) { break }
}
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
Write-Output ("RESULT card-off " + $cardKey + " after " + $t + " s, worker processes left " + (($w | Measure-Object).Count))
Start-Sleep -Seconds 5
}
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
# ---- the runs: --check for the build line and the self-test, --bench for the fingerprint and the rate ----
foreach ($pk in $packList) {
$d = Join-Path $packs $pk
Write-Output "RESULT run $pk check start $(Stamp)"
& $cuda --check --pack $d 2>&1 | ForEach-Object { "RESULT check $pk $_" }
Write-Output "RESULT run $pk check exit $LASTEXITCODE"
if ($hasBench) {
Write-Output "RESULT run $pk bench start $(Stamp)"
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT bench $pk $_" }
Write-Output "RESULT run $pk bench exit $LASTEXITCODE"
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT bench8 $pk $_" }
}
}
& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
# ---- restore the card with the settings it had ----
if ($cardOff) {
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
$en = $true; if ($null -ne $cardPref.enabled) { $en = [bool]$cardPref.enabled }
$body = @{ cards = @(@{ key = $cardKey; enabled = $en; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $cardKey + " enabled=" + $en) } catch { Write-Output ("RESULT error card restore " + $cardKey + ": " + $_.Exception.Message) }
}
Write-Output "RESULT job ca3-derive-pc2 end $(Stamp)"
exit 0