diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md index 1d47b2f7..4fca85cf 100644 --- a/docs/analysis/chip-model-v3.md +++ b/docs/analysis/chip-model-v3.md @@ -98,3 +98,46 @@ order: The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM. + +## 6. The per-day derivation (item 2) + +6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything +PROPOSED, a prototype behind load class `dr736`). The fixed-shape mixer of section 2's rows is replaced by nine +straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms, +the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item +derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program +of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a +32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's +constants in the wires. The counts are from the code (`memhard::mixer`: 144 ops per application as written, 128 +with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1; +the day program's floor is those counts, so the bare row cannot fall below x8's. + +| Row | Derivation | Chip ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) | Allowance 1.5x (cautious upper bound, approximate) | The old 3x (the fixed shape's; does not apply) | Equal silicon, SRAM deducted, at 1.2x / 1.5x | +|---|---|---|---|---|---|---|---|---|---|---| +| x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) | fixed mixer, 72 x 128 | 1,179,648 | 42.4 MH/s | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x | +| **dr736, the genesis day's draw** (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) | the day program, 9 x 736 instructions | 1,278,976 | 39.1 MH/s | 256 MiB | 128 / $46 | 0.29x | 0.34x | 0.43x | 0.86x | 0.29x / 0.36x | +| dr736 at the floor (a day whose draw sits exactly on the acceptance floor) | the day program | 1,179,648 | 42.4 | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x | +| dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) | the day program, 9 x 368 | 640,512 | 78.1 | 256 MiB | 128 / $46 | 0.57x | 0.69x | 0.86x | 1.72x | 0.57x / 0.71x | +| dr736 at year 4 (cache 512 MiB) | the day program | 1,278,976 | 39.1 | 512 MiB | 255 / $111 | 0.29x | 0.34x | 0.43x | 0.86x | 0.23x / 0.28x | + +Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2 += 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357. +The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or +divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in; +with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function +is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector, +and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional +compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious +upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule +(every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no +intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave +items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument, +history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's +2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged; +the 5090's build and compile are the PC 2 job, the 9070 XT's OWED. + +What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their +margins apply); no chip has been priced for its instruction store or its register files per item in flight; the +random ARX programs have had no cryptanalysis (the item 3 brief should name them beside `M_r`); the 2019-class +core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0 +carries that verdict. diff --git a/docs/bench-log.md b/docs/bench-log.md index 85997dca..53a6b8e3 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1961,3 +1961,84 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU. Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan. + +## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation + +Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and +the chip row in `docs/plans/counter-asic-3-derivation.md` and `docs/analysis/chip-model-v3.md` section 6. +Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x +to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash +idea, `superscalar.cpp` read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op +count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build +and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average. + +The construction (class `dr736`, `igneum-pow/src/derive.rs`): nine straight-line programs of 736 instructions per +item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64 +stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by `c|1`, rotate-add, xor-rotate, +add-constant, xor-constant, the `M_r` form `(d ^ c) * odd`, `d * odd + c`, `d ^= c & b`, `d += c | b`), every +instruction reading the register the previous one wrote (the chain, `s[0]` first) and writing another, every form +a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations, +or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written, +1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624 +instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major +interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT. + +**Verifier per 32-lane unit, one core (`with-lock.sh measure`, one session 07:42:20 to 07:42:33 UTC, load +average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; `igneum-pow bench --seed igneum-genesis +--day 2026-10-03 --class --warps 50`, two rounds, then the devnet seeds once):** + +| Class | ms per unit, avg of 50 (round 1 / 2) | Worst cold of three | Against x8 | Ops per item (GPU / chip / mul) | +|---|---|---|---|---| +| v2 | 0.598 / 0.594 | 0.697 | | 1,296 / 1,152 / 144 | +| x8 (mx8, class v3) | 2.061 / 2.063 | 2.179 | 1 | 10,368 / 9,216 / 1,152 | +| dr736 | 4.875 / 4.944 | 5.241 | 2.37x | 10,659 / 9,992 / 1,461 | +| x8, devnet seeds | 2.078 | 2.155 | | | +| dr736, devnet seeds | 4.872 | 5.241 | 2.34x | 10,701 / 10,083 / 1,362 | +| dr368 (half length, the x4-equivalent fallback) | 2.692 | 2.898 | 1.31x | 5,350 / 5,004 / 752 | + +The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core +figures. The interpreter's cost split (`examples/derive_perf.rs`, a functional run under the run lock, load 3.9 to +4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair +dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the +vector body (NEON, 1,180 `.4s` instructions in the binary). + +**Bit-exactness (`with-lock.sh run`):** dr736-genesis on Metal (`packbench --batches 1 --batch-log2 24`) cache +FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint +2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (`igneum-bench-cl-dr736-genesis --bench-pack`) the +self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0 +on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset +and on 2^24 outputs. + +**Daily 1 GiB build and hash rate, Metal (`with-lock.sh measure`, the same session, `packbench --batches 2 +--batch-log2 22 --group 256`, three rounds):** + +| Pack | Compile (1 / 2 / 3) | Build, GPU ms (1 / 2 / 3) | MH/s GPU (1 / 2 / 3) | +|---|---|---|---| +| mx8-genesis (x8, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | +| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | + +**Chip model (chip-model-v3.md section 6):** 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at +50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no +longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a +cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x. + +**Consequences per tier.** The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms +(x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs +11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core +(2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The +build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below; +the 9070 XT is OWED (PC 1 is the project lead's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier +already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to +94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the +same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; +the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line +is the number to read. The hash rate: unchanged within 0.3% on the Mac, as the hash kernel only loads. Packs grow +by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier. + +**Go / no-go:** GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in +docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms; +the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate +laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare. + +**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2 +lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED. diff --git a/docs/plans/counter-asic-3-derivation.md b/docs/plans/counter-asic-3-derivation.md new file mode 100644 index 00000000..902043af --- /dev/null +++ b/docs/plans/counter-asic-3-derivation.md @@ -0,0 +1,415 @@ +# Counter ASIC 3.0 item 2: a random item-derivation program per day + +6 October 2026. Worker `derive` (branch `ca3-derive`), under the brief of `docs/plans/counter-asic-3.md` item 2 and +the history audit's addition 2 (`docs/analysis/asic-resistance-history.md` section 4.3). Everything here is +PROPOSED: a prototype behind a load class (`dr736`), measured on the Mac and on PC 2, written as a reserve entry for +spec 1.13.2 (section 6) that lives in this file until the project lead's word. Nothing is published and no vector of class v2 +or v3 moves (section 3.4). + +What it does, in one line: the fixed-shape mixer `M_r` of spec 1.8.4, applied 72 times per item under class v3, +is replaced by nine straight-line programs of 736 instructions drawn once a day from the day key stream, so the chip +that holds the cache on die must run an arbitrary program instead of a wired pipeline; the 8 dependent cache reads +per item, the cache, the loads and the hash kernel are untouched. + +## 0. The verdict first + +| Gate | Number | Bar | Result | +|---|---|---|---| +| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) | +| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes | +| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 | +| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is the project lead's desk today) | +| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change | +| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies | + +Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class +core measurement (O-1.14) lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core +(pass) against about 12 ms on the approximate laptop row (fail). The class that passes both rows today is `dr368` +(2.69 ms), at the x4-equivalent op count, with the chip row at 0.58x bare. The way to the 736 figure under the gate +on a laptop is the JIT (section 4.3), which is out of scope and named with its risk. + +## 1. Why: what the chip model says the fixed shape is worth + +`docs/analysis/chip-model-v3.md` section 2 prices the on-die-cache recompute chip at 50 T op/s: class v3 (x8) costs +it 1,198,080 integer ops per hash (72 mixer applications x 128 items x about 130 ops), 41.7 MH/s, 0.31x the 5090's +136.1 MH/s bare, and 0.92x with "the 3x fixed-function factor (approximate, from memory: 2x to 5x is the usual +credit for a pipeline with no scheduling or divergence)". That credit is the mixer's fixed shape: the chip unrolls +the 72 applications into a wired pipeline with the day's constants baked in, no instruction fetch, no register +file, no operand muxes. RandomX's answer (`vendor/RandomX/src/superscalar.cpp` is not in this tree; read on +6 October 2026 from the upstream repository at commit 7607fb2 into the session scratchpad; the design argument is +`doc/design.md`, history section 2.4) is SuperscalarHash: the dataset item derivation is itself a random program +drawn from the cache key, 8 programs of about 450 instructions per item, so a light-mode chip "becomes a CPU". +RandomX's generator (`generateSuperscalar`, lines 653 to 850) schedules for a superscalar x86 core: it picks a +decode-buffer configuration per cycle, selects a source register that is ready at the cycle and a destination +that is not the source and was not last written by the same op group (`selectDestination`, line 495: no +"xor r,r2; xor r,r2", no "ror r,C1; ror r,C2", no two multiplies in a row on one register), and then computes the +program's ASIC latency as the longest dependency chain (lines 810 to 824) and sets the address register to the +register with the highest one. The item init is `rl[0] = (itemNumber + 1) * superscalarMul0`, the other seven +registers `rl[0] ^ superscalarAdd_i`, then per cache access: run the program, XOR the mix block in, next address +from the address register (`dataset.cpp` `initDatasetItem`, lines 164 to 190). + +What carries over here and what does not: the per-day program, the acceptance by construction, the dependency +chain and the address register idea carry over; the x86 port scheduling does not (our verifier interprets the +program for 32 items at once and our miners compile it for a GPU, so the schedule that matters is the GPU's), and +the latency bound RandomX relies on (the program's critical path against DRAM) is replaced by ours: the 8 +dependent cache reads per item, which are untouched. + +## 2. The design + +### 2.1 The draw + +One SplitMix64 stream seeded with `K[0] | (K[1] << 32)` (the day key, spec 1.8.1), the 40 draws of 1.8.4 +(`ROT`, `MUL`, `RC`) first, exactly as today (the item init `s[8 + i] = t * MUL[i] + RC[i]` still uses them), then +the program: `DERIVE_PROGRAMS = 9` round programs (one before each of the 8 cache reads, one after the last) of +`derive_len` instructions each, four draws per instruction in a fixed order: + +| Draw | Range | Sets | +|---|---|---| +| `below(100)` | op roll | the form, by the weight table of 2.2 (cumulative) | +| `below(15)` | destination roll | `d`: the 15 registers other than the chain `c`, in ascending order (roll >= c adds one) | +| `below(14)` | third-register roll | `b`: the 14 registers other than `d` and `c`, ascending (two skips); used by `andx` and `orx`, consumed by every form | +| `next()` | the immediate | `k = 1 + (low32 mod 31)` for the rotate forms; `imm = low32` for `addc`, `xorc`; `imm = low32 OR 1` for `mulc`, `mulc2`; consumed by every form | + +The chain register `c` is `s[0]` (the address word) at the start of each round program and the destination of the +previous instruction after it. Every instruction reads `c` and writes `d != c`; the next instruction's chain is +`d`. So no two instructions of a program can run in parallel (the "every instruction consumes the newest result" +rule of the brief, SuperscalarHash's chain made strict), and no two consecutive instructions write one register, +which is what removes the mergeable pairs SuperscalarHash's `selectDestination` guards against. Four draws per +instruction whether the form uses them or not, so the stream position of every draw is fixed by its index and a +future change to one form's draw changes no other draw (the rule 1.13.1 follows for `epoch_len`). A rejected +candidate (2.4) is followed by the next: the stream continues, `attempt + 1`, as the program generator of 1.4.6. + +Source: `igneum-pow/src/derive.rs` (`DeriveProgram::draw_candidate`, `draw`), `memhard.rs` (`MixParams::with_shape`). + +### 2.2 The instruction set: twelve two-register forms, fixed at genesis + +`c` the chain, `d` the destination, `b` the third register, `k` in 1..31, `i` a 32-bit constant (odd for the +multiplies). All arithmetic modulo 2^32, no division, no float, no data-dependent branch. Every form is a bijection +on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply, or a rotation of itself; `c` +and `b` are not written), so a program loses no entropy, the property `M_r` has. + +| Form | Semantics | Weight (percent) | GPU ops | Chip ops | Multiply | +|---|---|---|---|---|---| +| `add` | `d += c` | 14 | 1 | 1 | | +| `sub` | `d -= c` | 10 | 1 | 1 | | +| `xor` | `d ^= c` | 14 | 1 | 1 | | +| `mul` | `d *= (c OR 1)` | 10 | 2 | 1 (the OR is a wire) | yes | +| `rot` | `d = rotl(d, k) + c` | 10 | 2 | 2 | | +| `xrot` | `d = rotl(d ^ c, k)` | 10 | 2 | 2 | | +| `addc` | `d += c + i` | 6 | 2 | 2 | | +| `xorc` | `d ^= c ^ i` | 6 | 2 | 2 | | +| `mulc` | `d = (d ^ c) * i` (the per-word form of `M_r`, the chain in place of the round constant) | 8 | 2 | 2 | yes | +| `mulc2` | `d = d * i + c` | 4 | 2 | 2 | yes | +| `andx` | `d ^= (c AND b)` | 4 | 2 | 2 | | +| `orx` | `d += (c OR b)` | 4 | 2 | 2 | | + +Mean 1.62 GPU ops and 1.52 chip ops per instruction, 22% multiplies. `and` and `or` enter only as `andx` and +`orx` (a destructive `d &= c` would lose bits; the XOR and add of a conjunction keep `d` invertible). The forms are +the brief's set (add, sub, mul-lo, xor, rotate by 1..31, and, or, the ARX-multiply forms of `M_r`); `mulhi` is in +the lottery hash's families and bit-exact on the three vendors, and is left out of the derivation on purpose so +every form is one that the three compilers lower to a single integer instruction (section 3.3). + +### 2.3 Length and the floors: the x8-equivalent op count + +The coordinator's rule for item 2: the total op count per item equals or exceeds today's x8 count, so the chip's +budget row does not fall. The x8 mixer counted from the code (`memhard::mixer`): 16 x (xor, add, mul) + 8 quarter +rounds x 12 = 144 ops as written, 128 with the `RC[i] + rk` adds hoisted as constants (the chip and every compiler +do that), 16 multiplies; 72 applications per item = 10,368 ops as written, 9,216 hoisted, 1,152 multiplies +(`chip-model-v3.md` section 1 prices 130 per application from the spec text; the item 3 worker counted the same +144 / 128, coordinator's note of 6 October). The floors are those three. `derive_len = 736` gives 6,624 +instructions per item, expected 10,731 GPU ops, 10,068 chip ops and 1,457 multiplies (the genesis day draws +10,659 / 9,992 / 1,461; the devnet day 10,701 / 10,083 / 1,362), 7 to 22 standard deviations above the floors, so a +rejection on a floor is a rare event and the test exists for the degenerate class. The class name is `dr736`; the +half-length `dr368` (the x4-equivalent) is measured beside it as the fallback with its floors scaled. + +### 2.4 The acceptance test + +`DeriveProgram::check`, on every candidate; a rejection draws the next attempt from the stream: + +| Test | Rejects | Expected rate at 736 | +|---|---|---| +| every register written in every round program | a register no instruction of a round program writes (its init word would never enter the chain within that round) | 16 x 9 x (14/15)^736 = under 10^-20 | +| at least 8 distinct rotation amounts across the item's programs | all rotations equal (the brief's degenerate draw; the mixer's "ROT draw of eight equal values is possible and untested", spec 1.8.4) | about 1,300 rotate forms over 31 values: never | +| chip ops >= 9,216, GPU ops >= 10,368, multiplies >= 1,152 per item (scaled to the length) | an op draw under the x8 count | 22, 9 and 9 standard deviations below the mean: never | +| structural (asserted, hold by construction): `d != c`, `b` distinct from both, `k` in 1..31, odd multiplier constants | a generator bug | n/a | + +The generator panics after 64 rejected candidates in a row (`MAX_ATTEMPTS`), which the rates above put beyond +any day the chain will see; the panic is the right failure (a node that cannot derive the day's program cannot +verify, and must say so rather than guess). The weights are fixed at genesis; only the order, the registers and +the constants are drawn, so the family mix of a program cannot be steered by the draw. + +### 2.5 The item, with the program in place + +Spec 1.8.5 under the derivation class (`derive_len` nonzero, `mixer_mult` unused): + +``` +s[i] = K[i] for i in 0..7 +s[8 + i] = t * MUL[i] + RC[i] for i in 0..7 +for r in 0..7: + s = P_r(s) round program r, the chain starting at s[0] + a = s[0] AND (2^(C - 4) - 1) cache line index, as today + s[i] = s[i] XOR cache[line a][i] for i in 0..15 +s = P_8(s) +item(t) = s +``` + +The 8 dependent cache reads per item are exactly today's: the address of read `r` is `s[0]` after program `r`, and +`s[0]` depends on every earlier read through the chain (every round program writes every register, 2.4, and the +XOR of the line into all 16 words feeds the next program). The verifier's latency part (8 dependent misses per +item, overlapped across the up to 32 items of a load, spec 1.11) is unchanged, which the measurement shows: the +difference against x8 is the ALU part only (section 5.1). + +## 3. The prototype + +### 3.1 Where it lives + +| Item | Where | +|---|---| +| `DOp`, `DInstr`, `DeriveProgram` (draw, check, counts, fingerprint), the SoA interpreter `run_round` (pair dispatch), the scalar reference `run_round_scalar`, the text forms `instr_text` and `instr_line` | `igneum-pow/src/derive.rs` | +| `Shape::derive_len`, `Shape::is_derived`, `MixParams::derive` (drawn after the 40 mixer draws), `derive_items` dispatching to `derive_items_program` | `igneum-pow/src/memhard.rs` | +| `LoadClass::derive_len`, `LoadClass::DR736`, `with_derive`, parse and name `dr`, the program id (`derive/` + the length) | `igneum-pow/src/generator.rs` | +| `mh_round_0..8` and the program-driven `mh_item` in memhard.h, memhard.metal and kernel.cl; `IGNEUM_DERIVE_*` in program.h; `derive_len`, the op mix, the floors' counts and the nine programs (one line per instruction) in program.json | `igneum-pow/src/emit.rs` | +| `--class dr736` (or any `dr`) on every command; the bench prints the program's counts | `igneum-pow/src/main.rs` | +| `examples/derive_perf.rs`: the interpreter's cost per instruction per batch, drawn program against uniform programs | `igneum-pow/examples/` | +| Packs `dr736-genesis` (seed igneum-genesis, day 2026-10-03, program id 72c1d8048aef9542, program fingerprint 463535d01511350d) and `dr736-devnet-epoch0` (the devnet epoch 0 and day seeds, 7f4a5ca0a3637820, 771868df4e64d6ab); generator 2 with the class in the id, the x4-record shape, not a class v3 pack | `proto-cuda/packs-ca3-derive/` | +| Tests: `src/derive.rs` (5), `tests/derive.rs` (7: by hand on a small cache, batches, v2 and v3 untouched, the stream and the class, determinism and the pack text, stats beside x8, the text forms against the scalar reference, the word path) | `igneum-pow` | +| The PC 2 job | `relay/playbooks/ca3-derive-pc2.ps1` (section 5.4) | + +### 3.2 The verifier without a JIT: the word-major interpreter + +The verifier derives up to 32 distinct items per load (one per lane of the unit, `MemhardCpu::fetch`). The +interpreter keeps the 32 item states word-major (`st[reg][lane]`, 2 KiB) and runs each instruction across the +whole batch in one straight loop the compiler vectorises (NEON `add.4s`, `mul.4s`, `ushl.4s` and so on: 1,180 +such instructions in the example binary), so the dispatch is paid once per instruction per batch, not per item; +the cache reads of the batch are issued together after each round program, as the fixed-mixer loop does, so the 8 +dependent misses of independent items overlap. The dispatch is on PAIRS of instructions (144 arms, one indirect +branch per two instructions): a drawn op sequence is random, the predictor misses most dispatches, and pairing +halves the misses per instruction. Measured with `examples/derive_perf.rs` (a functional run, load average 3.9 to +4.9): 10.7 ns per instruction per batch cold, 7.18 warm with single dispatch, 4.98 with pair dispatch; a uniform +program of one form (predictable dispatch) 3.4 to 4.3 ns, so the body is about 3.5 ns and the remaining dispatch +cost about 1.5 ns. 6,624 x 4.98 ns x 128 batches = 4.2 ms per unit of interpreter time; the measured 4.88 ms +includes the latency part and the transposes. + +### 3.3 Bit-exactness on three vendors, by construction + +Each form is one C statement on `uint` with `+`, `-`, `^`, `*`, `|`, `&` and the memhard core's `mh_rotl` (a +shift pair, `n` in 1..31 at every call site), the same text in Metal, CUDA C and OpenCL C, every operand a 32-bit +unsigned integer: the same argument as spec 1.14 for the lottery hash's families, which have run bit-exact on the +three vendors since 4 October. The program is emitted as nine functions of straight-line statements (196 KB of +memhard.h per day); NVRTC, the Metal compiler and the OpenCL compilers see no loop, no branch and no call inside a +round program. Measured: section 5.2 (Metal and Apple OpenCL), 5.4 (CUDA). + +### 3.4 The v2 and v3 paths are untouched + +`Shape::derive_len` is 0 and `LoadClass::derive_len` is 0 on `V2`, `MX4`, `MX8` and `V3_CLASS`; `MixParams::derive` +is `None`; `derive_items` takes the fixed-mixer loop as before; the emitter's text for a shape without a program is +unchanged. `cargo test -p igneum-pow` (commit acb96ee): 58 lib, 7 derive, 4 mixer, 19 packs (every pinned pack of +v2 and v3 regenerated and compared byte for byte), 7 scratch, all green. + +## 4. What the chip keeps, and the fallbacks + +### 4.1 The fixed-function allowance after the change + +The chip of `chip-model-v3.md` now has to execute, per item, 6,624 instructions from a 12-form set over a +16-entry register file, with three operand fields and a constant per instruction, in an order and with operands +that change every day. That is a sequencer: an instruction store (6,624 x 9 bytes = 60 KB per day program, in +SRAM beside the cache mirror), a register file with three read ports and one write port, a 32-bit ALU with a +multiplier and a barrel rotator, operand muxes, and a program counter. A GPU streaming multiprocessor is the same +machine with a wider register file and a warp scheduler. What the chip keeps over the GPU: no warp scheduler, no +operand collector, no instruction cache hierarchy for a 1 MB kernel (the GPU's day program is about 60,000 SASS +instructions per thread; whether the 5090's instruction cache holds it is in the build time of 5.4), no graphics +or float units idle on the die, and the day's constants folded into the instruction store. What it loses: the +wired pipeline (the mixer's 72 applications as 72 stages with the constants in the wires, no fetch, no register +file, no crossbar), which is the thing the 3x credit paid for. The chain rule adds a second loss: with no +intra-item parallelism a single engine finishes one instruction per cycle at best and its multiplies serialise +unless it interleaves items, which costs a register file per item in flight (RandomX's light-mode argument, +history 2.4: a chip paying "760 cycles and 1,240 multiplies per item"). ProgPoW claimed 1.1x to 1.2x for exactly +this kind of chip ("conventional compute chips gain little on ProgPoW", Bob Rao's hardware audit, history 2.4, +[S70]; EIP-1057's own claim 1.1x to 1.2x, [S67]). So the allowance this document carries is 1.2x (ProgPoW's +claimed range, cited) with 1.5x as the cautious upper bound (approximate, mine), against the 3x of the fixed shape +(approximate, from memory, M16). Section 7 prices all of 1.0x, 1.2x, 1.5x, 2x and 3x. + +### 4.2 What it does not change + +The partial-store chip (item 1): a chip that stores the dataset and never derives items pays nothing for the +program; item 1's rows stand on their own. The cryptanalysis question (item 3) changes shape: instead of one +fixed `M_r` to attack, the attacker gets a fresh random ARX program every day, which is RandomX's bet and is +untested here; the day's programs can be audited by the same tools as random ARX ciphers, and the acceptance test +is the place to add a structural rule if one is found. The era draws and the cache growth are untouched. + +### 4.3 The JIT, named as the fallback with its risk + +A per-day JIT for the verifier (emit NEON or AVX2 code for the 32-lane batch, the chain row kept in registers +across instructions since `c` is always the row just written) would remove the 1.5 ns dispatch and about a third +of the 3.5 ns body (the chain row's loads), about 2.5 ms per unit on this core, the x8 figure. Its risk: a code +generator in the consensus path on two architectures (arm64, x86-64) whose output must equal the interpreter's +bit for bit, writable-executable memory in a node and in every pool verifier, a new attack surface the history's +lesson 9 says to audit before launch, and a second implementation per platform to keep in lockstep. Out of scope +for this item; it is the route to 736 under the gate on a laptop core if the measurement of O-1.14 confirms the +approximate row, and `dr368` is the route that needs no JIT. + +## 5. Measurements + +All under the locks of the brief; every row says its lock and the load average. Machine: Apple M5 Max, 64 GiB, +Darwin 25.6.0. Commit acb96ee (the code and the packs), packbench built from this worktree. + +### 5.1 The verifier per unit, one M5 Max core (`with-lock.sh measure`, one session, 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end) + +Script `measure-ca3-derive.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03 +--class --warps 50`, two rounds, then the devnet seeds (`--epoch-hex edc4fa84...fb07 --day-hex +69676e65756d2d6461792ffa50000000000000`) and `dr368` once. + +| Class | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | Against x8 | Items per unit | Ops per item (GPU / chip / multiplies) | +|---|---|---|---|---|---|---| +| v2 (igneum-genesis) | 0.598 / 0.594 | 0.697 | 1 | | 4,096 | 1,296 / 1,152 / 144 (9 x 144, hoisted 9 x 128) | +| x8, mx8 (class v3) | 2.061 / 2.063 | 2.179 | 3.46x | 1 | 4,096 | 10,368 / 9,216 / 1,152 | +| **dr736** | **4.875 / 4.944** | **5.241** | 8.2x | 2.37x | 4,096 | 10,659 / 9,992 / 1,461 | +| x8, the devnet seeds | 2.078 | 2.155 | | | 4,095 to 4,096 | | +| dr736, the devnet seeds | 4.872 | 5.241 | | 2.34x | 4,096 | 10,701 / 10,083 / 1,362 | +| dr368 (the x4-equivalent fallback) | 2.692 | 2.898 | 4.5x | 1.31x | 4,096 | 5,350 / 5,004 / 752 | + +Reading: the v2 row reads 0.59 to 0.60, the quiet readwidth night's 0.604 to 0.626 and 6.4a's 0.607 to 0.611, so +this session is a quiet-core figure and no scaling applies. At the same chip-op count as x8 (plus 8%), the +interpreter costs 2.4x what the compiled mixer costs: 4.98 ns per instruction per batch (3.2), of which about +1.5 ns is dispatch and 3.5 ns the vector body, against the mixer's compiled straight-line loop over the same +batch. The latency part is the same in every row (the 8 dependent misses per item; the dr368 row at half the +instructions saves 2.2 ms of the 4.9, which puts the latency-and-transpose share at about 0.5 ms). The 10 ms gate +keeps 5.1 ms (worst cold 4.76 ms) at 736 on this core and 7.3 ms at 368. + +### 5.1a Verification throughput per tier (the form of mixer-x4.md 6.5) + +| Figure | x8 (this session) | dr736 | dr368 | Note | +|---|---|---|---|---| +| ms per unit, quiet M5 Max core (measured, this session) | 2.06 | 4.88 | 2.69 | | +| ms per unit, 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) | 5.2 | 12.2 | 6.7 | the figure that fixes the gate is a measurement, not this row | +| Shares per second per core (quiet M5 Max) | 485 | 205 | 372 | | +| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second) | 4.5 | 10.7 | 5.9 | a pool verifying two units at once would halve the dispatch share (3.2); a 64-lane interpreter is a follow-up, unmeasured | +| Node: worst cold single unit (per block) | 2.2 ms | 5.2 ms | 2.9 ms | a block's verification stays under the 1 s block time by 190x | +| IBD over 108,000 headers on one core | 3.7 min | 8.8 min | 4.8 min | laptop (approximate): 9.4 / 22 / 12 min | +| Margin left under the 10 ms gate (worst cold, this core) | 7.8 ms | 4.8 ms | 7.1 ms | on the laptop row (approximate): 4.8 / none (over by 2.2) / 3.3 ms | + +Consequences per tier: every miner tier is untouched by the verifier (the miner never runs it); a pool operator +pays 2.4x the cores at 736 (11 cores for a 22,000-member pool against 4.5) or 1.3x at 368; a node on any 2026 core +verifies a block in 5 ms; a node on a 2019-class laptop core is the open question (O-1.14), and at 736 the +approximate row says it misses the gate, so 736 does not go genesis-live on an approximation, and 368 passes it. + +### 5.2 Bit-exactness on the Mac (`with-lock.sh run`, 08:38 to 08:41 local) + +`packbench --pack --batches 1 --batch-log2 24 --group 256` (Metal, built from this worktree) and +`igneum-bench-cl-dr736-genesis --bench-pack --pack --batches 1 --batch-log2 24` (Apple OpenCL, `proto-opencl/ +build.sh` on the pack through a temporary link). Vectors are the Rust interpreter's. + +| Pack | Harness | Cache FNV-1a 64 | Dataset head, word [MASK], 64 samples | Vectors | Fingerprint 2^24 | Compile | 1 GiB build, GPU ms (run lock, indicative) | +|---|---|---|---|---|---|---|---| +| dr736-genesis | Metal | 48c4f5bf24166b2e PASS | head and last PASS (packbench checks no samples) | 3/3 standalone, 3/3 in batch | 50e3eaa779da4f1e | 784 ms (276 on the second process, 1 ms once the shader cache has it) | 38.2 / 38.3 | +| dr736-genesis | Apple OpenCL | PASS (head, last line, FNV) | head, word [268435455], 64 samples PASS | 96 of 96 lanes | 50e3eaa779da4f1e | (in the 385 ms prepare) | 57 ms wall | +| dr736-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 9553f6d5c667205a | 763 ms | 37.4 | + +Reading: the Rust interpreter, Metal and Apple OpenCL agree on the derived dataset (head, word [MASK], the 64 +samples through OpenCL), on every vector lane and on the 2^24-output fingerprint of dr736-genesis across both +compilers; the devnet-seed pack agrees on Metal. The Apple OpenCL compile of a 6,624-statement item function +is inside a 385 ms prepare; the Metal compile is 0.75 to 0.8 s cold per day (the memhard library is the day's, +compiled once a day, not per epoch) and 1 ms from the shader cache. + +### 5.3 The daily build and the hash rate on the M5 Max (`with-lock.sh measure`, the session of 5.1, three rounds) + +`packbench --pack --batches 2 --batch-log2 22 --group 256`, mx8-genesis then dr736-genesis, three rounds. + +| Pack | Compile (round 1 / 2 / 3) | 1 GiB build, GPU ms (round 1 / 2 / 3) | MH/s GPU (round 1 / 2 / 3) | Vectors, self-tests | +|---|---|---|---|---| +| mx8-genesis (class v3, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | 3/3 + 3/3, PASS | +| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | 3/3 + 3/3, PASS | + +Reading: the build is 29 ms against 22 (+32%, +7 ms): the Mac's build was latency-bound at x1, x4 and x8 (21 to +22 ms at every multiplier, mixer-x4.md 6.4) and the serial chain per thread now shows (no intra-item ILP for the +compiler to schedule), still 34x under the 1 s bar. The hash rate is the x8 rate within 0.3% (27.06 to 27.13 +against 27.08 to 27.16), as it must be: the hash kernel only loads. Consequences per tier: a daily build of 29 ms +on Apple silicon costs nothing to any tier; the 5090's figure is section 5.4, the 9070 XT's is OWED (PC 1 not +released today; its x8 build was 72 to 77 ms and arithmetic-bound at none of x1, x4, x8, so the chain's cost there +is the open number); the integrated tier is the one to watch (5.5). + +### 5.4 RTX 5090, PC 2 (one job, `relay/playbooks/ca3-derive-pc2.ps1`) + +PENDING at the time of writing: the job waits for `/tmp/igneum-devnet/pc2-ca3.clear` (the proving agent's +30-minute measurement) and the `pc2-ca3.lock`. The job downloads the packs zip itself (sha256 +aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676: dr736-genesis, dr736-devnet-epoch0, mx8-genesis, +v2-genesis-mh), runs the installed app's igneum-worker-cuda.exe (NVRTC compiles each pack's own text) with the +NVIDIA card off in the app only under test (its key from settings.json), `--check` for the `nvrtc .. cache .. +dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8. +The rows land in the bench-log entry when the closing report is read. + +### 5.5 The build per tier, with the integrated tier + +| Card | Build at x8 | Build with the day program | Source | +|---|---|---|---| +| M5 Max, Metal | 22.1 ms | 29.0 ms (+32%) | 5.3, measure lock | +| RTX 5090, CUDA | 23 ms | section 5.4 | the PC 2 job | +| RX 9070 XT, OpenCL | 72 to 77 ms | OWED (PC 1) | | +| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | about 55 to 94 s (approximate, mixer-x4.md 6.5: the iGPU's build is arithmetic-bound at x1 already, scaled x8 from 6.9 / 9.4 / 11.7 s) | about the same count of ops at a lower ILP: 55 to 120 s (approximate, unmeasured) | `docs/plans/epoch-length.md` 6.1 iGPU rows | +| gfx1036 beside WSL build jobs (PC 1) | about 7 to 17 min (approximate) | the same or worse (approximate) | epoch-length.md 6.1 | +| 8 GB-class discrete card (not owned, about a tenth of the 5090, approximate) | about 1 s | about 1 to 1.3 s (approximate) | scaled | + +Consequences: nothing changes for a discrete card of any size on any vendor (the build is under a second), the +Mac row measured; the integrated tier already misses the per-prepare rule at x8 and needs the per-day dataset +reuse in the workers (0.3.12) or a restart per epoch, and the day program makes that need the same, not larger in +kind; the per-day compile of the item function (0.75 s Metal; NVRTC on the 5090 in 5.4) lands once a day in the +worker's day-cache build, not per epoch, unless the hash kernel's compile includes memhard.h, which on the CUDA +worker it does (kernel_bound.cu includes it, and the variant race compiles 17 variants): the 5.4 job's nvrtc line +is the number for that, and if it is large the fix is to compile the item function once per day into its own +module. + +## 6. PROPOSED spec text for 1.13.2: reserve entry R0, `derive` (the per-day item-derivation program) + +Not written into `docs/spec`; it lives here until the project lead's word. Named R0, ahead of R1 (mm8), because the reserve +is to be ordered by chip-unfriendliness (counter-asic-3.md item 6; mm8 last) and a derivation program is the most +chip-unfriendly entry the reserve can hold: it removes the fixed-function allowance of the recompute chip rather +than adding a family that chip can license. + +> Reserve entry R0, `derive` (the per-day item-derivation program). Semantics: section 1.8.5 under +> `derive_len = 736`: the nine mixer slots of the item derivation (one before each of the 8 cache reads, one after +> the last) each run a straight-line program of 736 instructions drawn from the day key stream of 1.8.4 after its +> 40 draws, four draws per instruction (`below(100)` the form, `below(15)` the destination among the registers +> other than the chain, `below(14)` the third register among those other than the destination and the chain, +> `next()` the immediate: `1 + low32 mod 31` for the rotate forms, `low32` for `addc` and `xorc`, `low32 OR 1` for +> `mulc` and `mulc2`); the chain is `s[0]` at the start of each program and the previous destination after; the +> twelve forms and weights of `docs/plans/counter-asic-3-derivation.md` section 2.2, fixed; every form a +> bijection on the state; the 8 dependent cache reads, the mixer constants of the item init, the cache and the +> dataset mapping of 1.8.5 unchanged; `mixer_mult` unused under R0. Acceptance test, per candidate, the next +> attempt on rejection (the stream continues): every register written in every round program; at least 8 distinct +> rotation amounts; per item at least 9,216 operations with constants folded, 10,368 as written and 1,152 +> multiplies (the x8 mixer's counts from `memhard::mixer`: 72 x 128, 72 x 144, 72 x 16; the floors scale with the +> length). Edge vectors, each a hand-built item run on every vendor: item 0, item 1, item 2^28 - 1 and item 2^32 - 1 +> of the genesis day on a 2^16-word cache; a program whose first instruction is each of the twelve forms with +> `d = 15`, `c = 0`, `b = 14`, `k = 31`, `i = 0xffffffff` (odd for the multiplies) on the all-ones state and on the +> all-zero state (the wrap of every form); the day of the pinned pack `dr736-genesis` (program fingerprint +> 463535d01511350d, dataset head `vectors.json`, 2^24 fingerprint 50e3eaa779da4f1e) and of `dr736-devnet-epoch0` +> (771868df4e64d6ab, 9553f6d5c667205a). Unlock: at the start of era n = 2 (DAA 31,104,000), or earlier by the 90% +> signalling path of section 5.7, or at genesis if the verifier on a 2019-class core (O-1.14) reads under 10 ms per +> unit at 736, else at the length that does (368 measured at 2.69 ms on an M5 Max core); never by a release. The +> verifier procedure: the word-major interpreter of `igneum-pow/src/derive.rs` (32 item states per batch, pair +> dispatch), 4.88 ms per unit on one M5 Max core (section 5.1), no JIT; a JIT is the named fallback (4.3). Vendor +> paths: none needed; every form is a single 32-bit integer statement on all three compilers (3.3). + +## 7. The chip model row + +Written into `docs/analysis/chip-model-v3.md` section 6 in the form of its section 2. Ops per hash: 128 items x +9,992 chip ops (the genesis day's draw; the floor 9,216) = 1,278,976 (floor 1,179,648); chip rate at 50 T op/s = +39.1 MH/s (floor 42.4); bare against 136.1 MH/s = 0.287x (floor 0.31x, the x8 row's figure, as the floor is x8's +count). The allowance rows: 1.0x 0.29x; 1.2x (ProgPoW's claim) 0.34x; 1.5x (cautious upper bound, approximate) +0.43x; 2x 0.57x; 3x (the fixed shape's, which no longer applies) 0.86x. Equal silicon (x 0.829): 0.24 / 0.29 / +0.36 / 0.48 / 0.71. The x8 row read 0.92x at 3x and 0.76x at equal silicon; at the same 0.31x bare the day program +takes the chip from 0.92x to 0.34x to 0.43x, which is the margin the item was for. dr368 (the fallback): 639,488 +chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x, with less margin than x8 had at +3x and more than x4 had (1.84x). + +## 8. What is unverified or owed + +| Item | State | +|---|---| +| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s | +| RX 9070 XT (PC 1) | OWED: PC 1 is the project lead's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released | +| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it | +| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` | +| The integrated tier's build with the day program | approximate (5.5); the gfx1036 measurement is a PC 2 OpenCL job, not run today (the one PC 2 job carries the 5090) | +| A 64-lane interpreter for pools (two units per batch) | unimplemented; it would cut the dispatch share for pool verifiers only | +| The NVRTC cost of memhard.h inside the per-epoch hash kernel compile and the variant race | the PC 2 job's nvrtc line; the fix, if large, is one module per day for the item function | diff --git a/relay/playbooks/ca3-derive-pc2.ps1 b/relay/playbooks/ca3-derive-pc2.ps1 new file mode 100644 index 00000000..328e98d0 --- /dev/null +++ b/relay/playbooks/ca3-derive-pc2.ps1 @@ -0,0 +1,118 @@ +# Igneum run job: Counter ASIC 3.0 item 2, the per-day item-derivation program (docs/plans/counter-asic-3-derivation.md), +# on PC 2's RTX 5090 (machine 1ccfe586), 6 October 2026. ONE job carries everything (the PC 2 rule of the ca3 brief): +# it downloads the packs zip itself from the downloads host (sha256 checked), finds the installed app's +# igneum-worker-cuda.exe (NVRTC compiles each pack's own kernel text, so no new worker build is needed), switches the +# NVIDIA card off in the app ONLY while the packs run (the card key from the app's settings.json, never /api/state; +# restored after with the settings it had), and never quits, restarts or updates the installed app. The numbers this +# job is for: per pack, the worker's own `nvrtc .. cache .. dataset .. ms` line (the daily 1 GiB build on the 5090), +# the self-test (bit-exactness: cache FNV, dataset head, word [MASK], 64 samples, 96 vector lanes against the Rust CPU +# interpreter), the 2^24 fingerprint at base nonce 0 (against the Mac: dr736-genesis 50e3eaa779da4f1e, +# dr736-devnet-epoch0 9553f6d5c667205a, mx8-genesis 7c28cfb06c5c65a9, v2-genesis-mh 25f96e7dce90bd4e) and the hash rate. +# Packs: dr736-genesis, dr736-devnet-epoch0 (the derivation class), mx8-genesis (the x8 control), v2-genesis-mh (the v2 +# control). Every result line starts with RESULT. The placeholder __DL_BASE__ is substituted at publish time; the +# downloads token never enters the repository. +$ErrorActionPreference = 'Continue' +function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) } +function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } +# everything lives in this job's own folder (the wiped-jobs-folder rule of tools/ci/kit-path-check.sh: nothing of +# another job's is reached) +$work = Join-Path $env:IGNEUM_JOB_DIR 'ca3-derive' +New-Item -ItemType Directory -Force -Path $work | Out-Null +$zipUrl = '__DL_BASE__/igneum-ca3-derive-packs.zip' +$zipSha = 'aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676' +$zip = Join-Path $work 'igneum-ca3-derive-packs.zip' +Write-Output "RESULT job ca3-derive-pc2 start $(Stamp) machine $env:COMPUTERNAME" + +# ---- the packs: downloaded by this job, sha256 checked, expanded fresh ---- +try { + [Net.ServicePointManager]::SecurityProtocol = [Net.SecurityProtocolType]::Tls12 + Invoke-WebRequest -Uri $zipUrl -OutFile $zip -TimeoutSec 120 -Headers @{ 'Cache-Control' = 'no-cache' } +} catch { Write-Output ("RESULT error download: " + $_.Exception.Message); exit 2 } +$got = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower() +if ($got -ne $zipSha) { Write-Output "RESULT error zip sha256 $got expected $zipSha"; exit 2 } +Write-Output "RESULT zip sha256 $got size $((Get-Item $zip).Length) OK" +$packs = Join-Path $work 'packs-ca3-derive' +if (Test-Path $packs) { Remove-Item -Recurse -Force $packs } +Expand-Archive -Path $zip -DestinationPath $work -Force +$packList = @('v2-genesis-mh', 'mx8-genesis', 'dr736-genesis', 'dr736-devnet-epoch0') +foreach ($pk in $packList) { + $d = Join-Path $packs $pk + if (-not (Test-Path (Join-Path $d 'memhard.h'))) { Write-Output "RESULT error pack $pk missing after extract"; exit 2 } + Write-Output ("RESULT pack $pk memhard.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'memhard.h')).Hash.ToLower() + " program.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'program.h')).Hash.ToLower()) +} + +# ---- the worker: the installed app's igneum-worker-cuda.exe, run in place (its NVRTC DLLs sit beside it) ---- +$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1 +if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe'; exit 2 } +$cuda = Join-Path $inst 'igneum-worker-cuda.exe' +Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) nvrtc_dlls $((Get-ChildItem $inst -Filter 'nvrtc*.dll').Count)" +$help = (& $cuda --help 2>&1 | Out-String) +$hasBench = $help -match '--bench' +Write-Output "RESULT worker-cuda has --bench: $hasBench" +if (-not $hasBench) { Write-Output 'RESULT note the installed worker has no --bench; --check gives the build time and the self-test, no fingerprint and no hash rate' } + +# ---- the app: the NVIDIA card's key and settings from settings.json; switched off through POST api/cards only ---- +$appDir = $env:IGNEUM_APP_DIR +if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' } +$urlFile = Join-Path $appDir 'app.url' +$url = $null +if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() } +$sj = Join-Path $appDir 'settings.json' +$cardKey = $null; $cardPref = $null +if (Test-Path $sj) { + try { + $settings = Get-Content -LiteralPath $sj -Raw | ConvertFrom-Json + if ($settings.cards) { + foreach ($p in $settings.cards.PSObject.Properties) { if ($p.Name -like 'nvidia:*') { $cardKey = $p.Name; $cardPref = $p.Value; break } } + } + } catch { Say ("settings.json: " + $_.Exception.Message) } +} +if ($cardKey) { + Write-Output ("RESULT card " + $cardKey + " enabled=" + $cardPref.enabled + " identities=" + $cardPref.identities + " power_pct=" + $cardPref.power_pct + " (settings.json)") +} else { + Write-Output 'RESULT card none in settings.json (no nvidia:* entry); the app keeps mining on the card and the numbers carry that load' +} +$cardOff = $false +if ($cardKey -and $url) { + $ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities } + $pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct } + $body = @{ cards = @(@{ key = $cardKey; enabled = $false; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; $cardOff = $true; Say 'card off requested' } catch { Write-Output ("RESULT error api/cards off: " + $_.Exception.Message) } + # wait for the worker process to go, up to 90 s (the process list, not /api/state) + $t = 0 + while ($t -lt 90) { + Start-Sleep -Seconds 5; $t += 5 + $w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue + if (-not $w) { break } + } + $w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue + Write-Output ("RESULT card-off " + $cardKey + " after " + $t + " s, worker processes left " + (($w | Measure-Object).Count)) + Start-Sleep -Seconds 5 +} +& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" } + +# ---- the runs: --check for the build line and the self-test, --bench for the fingerprint and the rate ---- +foreach ($pk in $packList) { + $d = Join-Path $packs $pk + Write-Output "RESULT run $pk check start $(Stamp)" + & $cuda --check --pack $d 2>&1 | ForEach-Object { "RESULT check $pk $_" } + Write-Output "RESULT run $pk check exit $LASTEXITCODE" + if ($hasBench) { + Write-Output "RESULT run $pk bench start $(Stamp)" + & $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT bench $pk $_" } + Write-Output "RESULT run $pk bench exit $LASTEXITCODE" + & $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT bench8 $pk $_" } + } +} +& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" } + +# ---- restore the card with the settings it had ---- +if ($cardOff) { + $ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities } + $pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct } + $en = $true; if ($null -ne $cardPref.enabled) { $en = [bool]$cardPref.enabled } + $body = @{ cards = @(@{ key = $cardKey; enabled = $en; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5 + try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $cardKey + " enabled=" + $en) } catch { Write-Output ("RESULT error card restore " + $cardKey + ": " + $_.Exception.Message) } +} +Write-Output "RESULT job ca3-derive-pc2 end $(Stamp)" +exit 0