Counter ASIC 3.0 item 2: the design, the Mac measurements, the chip-model row and the PC 2 job
docs/plans/counter-asic-3-derivation.md (the design, the acceptance test, the interpreter, the allowance argument, the measurements, the PROPOSED reserve entry R0 for 1.13.2, what is owed), docs/analysis/chip-model-v3.md section 6 (the per-day derivation rows at 1.0x to 3x allowances), the bench-log entry, relay/playbooks/ca3-derive-pc2.ps1 (one PC 2 job: self-fetched packs zip, the installed worker through NVRTC, the card off only under test with its key from settings.json). Verifier 4.875 / 4.944 ms per unit on one M5 Max core under the measure lock against x8's 2.061 / 2.063; Metal build 29 ms against 22; hash rate equal; bit-exact on Metal and Apple OpenCL. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
50ebdd64d9
commit
bcc2db992e
4 changed files with 657 additions and 0 deletions
|
|
@ -98,3 +98,46 @@ order:
|
|||
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point
|
||||
under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had
|
||||
no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.
|
||||
|
||||
## 6. The per-day derivation (item 2)
|
||||
|
||||
6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything
|
||||
PROPOSED, a prototype behind load class `dr736`). The fixed-shape mixer of section 2's rows is replaced by nine
|
||||
straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms,
|
||||
the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item
|
||||
derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program
|
||||
of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a
|
||||
32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's
|
||||
constants in the wires. The counts are from the code (`memhard::mixer`: 144 ops per application as written, 128
|
||||
with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1;
|
||||
the day program's floor is those counts, so the bare row cannot fall below x8's.
|
||||
|
||||
| Row | Derivation | Chip ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) | Allowance 1.5x (cautious upper bound, approximate) | The old 3x (the fixed shape's; does not apply) | Equal silicon, SRAM deducted, at 1.2x / 1.5x |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) | fixed mixer, 72 x 128 | 1,179,648 | 42.4 MH/s | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
|
||||
| **dr736, the genesis day's draw** (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) | the day program, 9 x 736 instructions | 1,278,976 | 39.1 MH/s | 256 MiB | 128 / $46 | 0.29x | 0.34x | 0.43x | 0.86x | 0.29x / 0.36x |
|
||||
| dr736 at the floor (a day whose draw sits exactly on the acceptance floor) | the day program | 1,179,648 | 42.4 | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
|
||||
| dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) | the day program, 9 x 368 | 640,512 | 78.1 | 256 MiB | 128 / $46 | 0.57x | 0.69x | 0.86x | 1.72x | 0.57x / 0.71x |
|
||||
| dr736 at year 4 (cache 512 MiB) | the day program | 1,278,976 | 39.1 | 512 MiB | 255 / $111 | 0.29x | 0.34x | 0.43x | 0.86x | 0.23x / 0.28x |
|
||||
|
||||
Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2
|
||||
= 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357.
|
||||
The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or
|
||||
divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in;
|
||||
with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function
|
||||
is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector,
|
||||
and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional
|
||||
compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious
|
||||
upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule
|
||||
(every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no
|
||||
intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave
|
||||
items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument,
|
||||
history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's
|
||||
2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged;
|
||||
the 5090's build and compile are the PC 2 job, the 9070 XT's OWED.
|
||||
|
||||
What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their
|
||||
margins apply); no chip has been priced for its instruction store or its register files per item in flight; the
|
||||
random ARX programs have had no cryptanalysis (the item 3 brief should name them beside `M_r`); the 2019-class
|
||||
core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0
|
||||
carries that verdict.
|
||||
|
|
|
|||
|
|
@ -1961,3 +1961,84 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep
|
|||
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
|
||||
|
||||
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.
|
||||
|
||||
## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation
|
||||
|
||||
Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and
|
||||
the chip row in `docs/plans/counter-asic-3-derivation.md` and `docs/analysis/chip-model-v3.md` section 6.
|
||||
Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x
|
||||
to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash
|
||||
idea, `superscalar.cpp` read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op
|
||||
count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build
|
||||
and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average.
|
||||
|
||||
The construction (class `dr736`, `igneum-pow/src/derive.rs`): nine straight-line programs of 736 instructions per
|
||||
item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64
|
||||
stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by `c|1`, rotate-add, xor-rotate,
|
||||
add-constant, xor-constant, the `M_r` form `(d ^ c) * odd`, `d * odd + c`, `d ^= c & b`, `d += c | b`), every
|
||||
instruction reading the register the previous one wrote (the chain, `s[0]` first) and writing another, every form
|
||||
a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations,
|
||||
or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written,
|
||||
1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624
|
||||
instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major
|
||||
interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT.
|
||||
|
||||
**Verifier per 32-lane unit, one core (`with-lock.sh measure`, one session 07:42:20 to 07:42:33 UTC, load
|
||||
average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; `igneum-pow bench --seed igneum-genesis
|
||||
--day 2026-10-03 --class <c> --warps 50`, two rounds, then the devnet seeds once):**
|
||||
|
||||
| Class | ms per unit, avg of 50 (round 1 / 2) | Worst cold of three | Against x8 | Ops per item (GPU / chip / mul) |
|
||||
|---|---|---|---|---|
|
||||
| v2 | 0.598 / 0.594 | 0.697 | | 1,296 / 1,152 / 144 |
|
||||
| x8 (mx8, class v3) | 2.061 / 2.063 | 2.179 | 1 | 10,368 / 9,216 / 1,152 |
|
||||
| dr736 | 4.875 / 4.944 | 5.241 | 2.37x | 10,659 / 9,992 / 1,461 |
|
||||
| x8, devnet seeds | 2.078 | 2.155 | | |
|
||||
| dr736, devnet seeds | 4.872 | 5.241 | 2.34x | 10,701 / 10,083 / 1,362 |
|
||||
| dr368 (half length, the x4-equivalent fallback) | 2.692 | 2.898 | 1.31x | 5,350 / 5,004 / 752 |
|
||||
|
||||
The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core
|
||||
figures. The interpreter's cost split (`examples/derive_perf.rs`, a functional run under the run lock, load 3.9 to
|
||||
4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair
|
||||
dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the
|
||||
vector body (NEON, 1,180 `.4s` instructions in the binary).
|
||||
|
||||
**Bit-exactness (`with-lock.sh run`):** dr736-genesis on Metal (`packbench --batches 1 --batch-log2 24`) cache
|
||||
FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint
|
||||
2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (`igneum-bench-cl-dr736-genesis --bench-pack`) the
|
||||
self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0
|
||||
on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset
|
||||
and on 2^24 outputs.
|
||||
|
||||
**Daily 1 GiB build and hash rate, Metal (`with-lock.sh measure`, the same session, `packbench --batches 2
|
||||
--batch-log2 22 --group 256`, three rounds):**
|
||||
|
||||
| Pack | Compile (1 / 2 / 3) | Build, GPU ms (1 / 2 / 3) | MH/s GPU (1 / 2 / 3) |
|
||||
|---|---|---|---|
|
||||
| mx8-genesis (x8, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 |
|
||||
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 |
|
||||
|
||||
**Chip model (chip-model-v3.md section 6):** 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at
|
||||
50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no
|
||||
longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a
|
||||
cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.
|
||||
|
||||
**Consequences per tier.** The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms
|
||||
(x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs
|
||||
11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core
|
||||
(2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The
|
||||
build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below;
|
||||
the 9070 XT is OWED (PC 1 is the project lead's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier
|
||||
already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to
|
||||
94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the
|
||||
same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache;
|
||||
the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line
|
||||
is the number to read. The hash rate: unchanged within 0.3% on the Mac, as the hash kernel only loads. Packs grow
|
||||
by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.
|
||||
|
||||
**Go / no-go:** GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in
|
||||
docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms;
|
||||
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
|
||||
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
|
||||
|
||||
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
|
||||
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.
|
||||
|
|
|
|||
415
docs/plans/counter-asic-3-derivation.md
Normal file
415
docs/plans/counter-asic-3-derivation.md
Normal file
|
|
@ -0,0 +1,415 @@
|
|||
# Counter ASIC 3.0 item 2: a random item-derivation program per day
|
||||
|
||||
6 October 2026. Worker `derive` (branch `ca3-derive`), under the brief of `docs/plans/counter-asic-3.md` item 2 and
|
||||
the history audit's addition 2 (`docs/analysis/asic-resistance-history.md` section 4.3). Everything here is
|
||||
PROPOSED: a prototype behind a load class (`dr736`), measured on the Mac and on PC 2, written as a reserve entry for
|
||||
spec 1.13.2 (section 6) that lives in this file until the project lead's word. Nothing is published and no vector of class v2
|
||||
or v3 moves (section 3.4).
|
||||
|
||||
What it does, in one line: the fixed-shape mixer `M_r` of spec 1.8.4, applied 72 times per item under class v3,
|
||||
is replaced by nine straight-line programs of 736 instructions drawn once a day from the day key stream, so the chip
|
||||
that holds the cache on die must run an arbitrary program instead of a wired pipeline; the 8 dependent cache reads
|
||||
per item, the cache, the loads and the hash kernel are untouched.
|
||||
|
||||
## 0. The verdict first
|
||||
|
||||
| Gate | Number | Bar | Result |
|
||||
|---|---|---|---|
|
||||
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
|
||||
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
|
||||
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
|
||||
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is the project lead's desk today) |
|
||||
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
|
||||
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
|
||||
|
||||
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
|
||||
core measurement (O-1.14) lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core
|
||||
(pass) against about 12 ms on the approximate laptop row (fail). The class that passes both rows today is `dr368`
|
||||
(2.69 ms), at the x4-equivalent op count, with the chip row at 0.58x bare. The way to the 736 figure under the gate
|
||||
on a laptop is the JIT (section 4.3), which is out of scope and named with its risk.
|
||||
|
||||
## 1. Why: what the chip model says the fixed shape is worth
|
||||
|
||||
`docs/analysis/chip-model-v3.md` section 2 prices the on-die-cache recompute chip at 50 T op/s: class v3 (x8) costs
|
||||
it 1,198,080 integer ops per hash (72 mixer applications x 128 items x about 130 ops), 41.7 MH/s, 0.31x the 5090's
|
||||
136.1 MH/s bare, and 0.92x with "the 3x fixed-function factor (approximate, from memory: 2x to 5x is the usual
|
||||
credit for a pipeline with no scheduling or divergence)". That credit is the mixer's fixed shape: the chip unrolls
|
||||
the 72 applications into a wired pipeline with the day's constants baked in, no instruction fetch, no register
|
||||
file, no operand muxes. RandomX's answer (`vendor/RandomX/src/superscalar.cpp` is not in this tree; read on
|
||||
6 October 2026 from the upstream repository at commit 7607fb2 into the session scratchpad; the design argument is
|
||||
`doc/design.md`, history section 2.4) is SuperscalarHash: the dataset item derivation is itself a random program
|
||||
drawn from the cache key, 8 programs of about 450 instructions per item, so a light-mode chip "becomes a CPU".
|
||||
RandomX's generator (`generateSuperscalar`, lines 653 to 850) schedules for a superscalar x86 core: it picks a
|
||||
decode-buffer configuration per cycle, selects a source register that is ready at the cycle and a destination
|
||||
that is not the source and was not last written by the same op group (`selectDestination`, line 495: no
|
||||
"xor r,r2; xor r,r2", no "ror r,C1; ror r,C2", no two multiplies in a row on one register), and then computes the
|
||||
program's ASIC latency as the longest dependency chain (lines 810 to 824) and sets the address register to the
|
||||
register with the highest one. The item init is `rl[0] = (itemNumber + 1) * superscalarMul0`, the other seven
|
||||
registers `rl[0] ^ superscalarAdd_i`, then per cache access: run the program, XOR the mix block in, next address
|
||||
from the address register (`dataset.cpp` `initDatasetItem`, lines 164 to 190).
|
||||
|
||||
What carries over here and what does not: the per-day program, the acceptance by construction, the dependency
|
||||
chain and the address register idea carry over; the x86 port scheduling does not (our verifier interprets the
|
||||
program for 32 items at once and our miners compile it for a GPU, so the schedule that matters is the GPU's), and
|
||||
the latency bound RandomX relies on (the program's critical path against DRAM) is replaced by ours: the 8
|
||||
dependent cache reads per item, which are untouched.
|
||||
|
||||
## 2. The design
|
||||
|
||||
### 2.1 The draw
|
||||
|
||||
One SplitMix64 stream seeded with `K[0] | (K[1] << 32)` (the day key, spec 1.8.1), the 40 draws of 1.8.4
|
||||
(`ROT`, `MUL`, `RC`) first, exactly as today (the item init `s[8 + i] = t * MUL[i] + RC[i]` still uses them), then
|
||||
the program: `DERIVE_PROGRAMS = 9` round programs (one before each of the 8 cache reads, one after the last) of
|
||||
`derive_len` instructions each, four draws per instruction in a fixed order:
|
||||
|
||||
| Draw | Range | Sets |
|
||||
|---|---|---|
|
||||
| `below(100)` | op roll | the form, by the weight table of 2.2 (cumulative) |
|
||||
| `below(15)` | destination roll | `d`: the 15 registers other than the chain `c`, in ascending order (roll >= c adds one) |
|
||||
| `below(14)` | third-register roll | `b`: the 14 registers other than `d` and `c`, ascending (two skips); used by `andx` and `orx`, consumed by every form |
|
||||
| `next()` | the immediate | `k = 1 + (low32 mod 31)` for the rotate forms; `imm = low32` for `addc`, `xorc`; `imm = low32 OR 1` for `mulc`, `mulc2`; consumed by every form |
|
||||
|
||||
The chain register `c` is `s[0]` (the address word) at the start of each round program and the destination of the
|
||||
previous instruction after it. Every instruction reads `c` and writes `d != c`; the next instruction's chain is
|
||||
`d`. So no two instructions of a program can run in parallel (the "every instruction consumes the newest result"
|
||||
rule of the brief, SuperscalarHash's chain made strict), and no two consecutive instructions write one register,
|
||||
which is what removes the mergeable pairs SuperscalarHash's `selectDestination` guards against. Four draws per
|
||||
instruction whether the form uses them or not, so the stream position of every draw is fixed by its index and a
|
||||
future change to one form's draw changes no other draw (the rule 1.13.1 follows for `epoch_len`). A rejected
|
||||
candidate (2.4) is followed by the next: the stream continues, `attempt + 1`, as the program generator of 1.4.6.
|
||||
|
||||
Source: `igneum-pow/src/derive.rs` (`DeriveProgram::draw_candidate`, `draw`), `memhard.rs` (`MixParams::with_shape`).
|
||||
|
||||
### 2.2 The instruction set: twelve two-register forms, fixed at genesis
|
||||
|
||||
`c` the chain, `d` the destination, `b` the third register, `k` in 1..31, `i` a 32-bit constant (odd for the
|
||||
multiplies). All arithmetic modulo 2^32, no division, no float, no data-dependent branch. Every form is a bijection
|
||||
on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply, or a rotation of itself; `c`
|
||||
and `b` are not written), so a program loses no entropy, the property `M_r` has.
|
||||
|
||||
| Form | Semantics | Weight (percent) | GPU ops | Chip ops | Multiply |
|
||||
|---|---|---|---|---|---|
|
||||
| `add` | `d += c` | 14 | 1 | 1 | |
|
||||
| `sub` | `d -= c` | 10 | 1 | 1 | |
|
||||
| `xor` | `d ^= c` | 14 | 1 | 1 | |
|
||||
| `mul` | `d *= (c OR 1)` | 10 | 2 | 1 (the OR is a wire) | yes |
|
||||
| `rot` | `d = rotl(d, k) + c` | 10 | 2 | 2 | |
|
||||
| `xrot` | `d = rotl(d ^ c, k)` | 10 | 2 | 2 | |
|
||||
| `addc` | `d += c + i` | 6 | 2 | 2 | |
|
||||
| `xorc` | `d ^= c ^ i` | 6 | 2 | 2 | |
|
||||
| `mulc` | `d = (d ^ c) * i` (the per-word form of `M_r`, the chain in place of the round constant) | 8 | 2 | 2 | yes |
|
||||
| `mulc2` | `d = d * i + c` | 4 | 2 | 2 | yes |
|
||||
| `andx` | `d ^= (c AND b)` | 4 | 2 | 2 | |
|
||||
| `orx` | `d += (c OR b)` | 4 | 2 | 2 | |
|
||||
|
||||
Mean 1.62 GPU ops and 1.52 chip ops per instruction, 22% multiplies. `and` and `or` enter only as `andx` and
|
||||
`orx` (a destructive `d &= c` would lose bits; the XOR and add of a conjunction keep `d` invertible). The forms are
|
||||
the brief's set (add, sub, mul-lo, xor, rotate by 1..31, and, or, the ARX-multiply forms of `M_r`); `mulhi` is in
|
||||
the lottery hash's families and bit-exact on the three vendors, and is left out of the derivation on purpose so
|
||||
every form is one that the three compilers lower to a single integer instruction (section 3.3).
|
||||
|
||||
### 2.3 Length and the floors: the x8-equivalent op count
|
||||
|
||||
The coordinator's rule for item 2: the total op count per item equals or exceeds today's x8 count, so the chip's
|
||||
budget row does not fall. The x8 mixer counted from the code (`memhard::mixer`): 16 x (xor, add, mul) + 8 quarter
|
||||
rounds x 12 = 144 ops as written, 128 with the `RC[i] + rk` adds hoisted as constants (the chip and every compiler
|
||||
do that), 16 multiplies; 72 applications per item = 10,368 ops as written, 9,216 hoisted, 1,152 multiplies
|
||||
(`chip-model-v3.md` section 1 prices 130 per application from the spec text; the item 3 worker counted the same
|
||||
144 / 128, coordinator's note of 6 October). The floors are those three. `derive_len = 736` gives 6,624
|
||||
instructions per item, expected 10,731 GPU ops, 10,068 chip ops and 1,457 multiplies (the genesis day draws
|
||||
10,659 / 9,992 / 1,461; the devnet day 10,701 / 10,083 / 1,362), 7 to 22 standard deviations above the floors, so a
|
||||
rejection on a floor is a rare event and the test exists for the degenerate class. The class name is `dr736`; the
|
||||
half-length `dr368` (the x4-equivalent) is measured beside it as the fallback with its floors scaled.
|
||||
|
||||
### 2.4 The acceptance test
|
||||
|
||||
`DeriveProgram::check`, on every candidate; a rejection draws the next attempt from the stream:
|
||||
|
||||
| Test | Rejects | Expected rate at 736 |
|
||||
|---|---|---|
|
||||
| every register written in every round program | a register no instruction of a round program writes (its init word would never enter the chain within that round) | 16 x 9 x (14/15)^736 = under 10^-20 |
|
||||
| at least 8 distinct rotation amounts across the item's programs | all rotations equal (the brief's degenerate draw; the mixer's "ROT draw of eight equal values is possible and untested", spec 1.8.4) | about 1,300 rotate forms over 31 values: never |
|
||||
| chip ops >= 9,216, GPU ops >= 10,368, multiplies >= 1,152 per item (scaled to the length) | an op draw under the x8 count | 22, 9 and 9 standard deviations below the mean: never |
|
||||
| structural (asserted, hold by construction): `d != c`, `b` distinct from both, `k` in 1..31, odd multiplier constants | a generator bug | n/a |
|
||||
|
||||
The generator panics after 64 rejected candidates in a row (`MAX_ATTEMPTS`), which the rates above put beyond
|
||||
any day the chain will see; the panic is the right failure (a node that cannot derive the day's program cannot
|
||||
verify, and must say so rather than guess). The weights are fixed at genesis; only the order, the registers and
|
||||
the constants are drawn, so the family mix of a program cannot be steered by the draw.
|
||||
|
||||
### 2.5 The item, with the program in place
|
||||
|
||||
Spec 1.8.5 under the derivation class (`derive_len` nonzero, `mixer_mult` unused):
|
||||
|
||||
```
|
||||
s[i] = K[i] for i in 0..7
|
||||
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
|
||||
for r in 0..7:
|
||||
s = P_r(s) round program r, the chain starting at s[0]
|
||||
a = s[0] AND (2^(C - 4) - 1) cache line index, as today
|
||||
s[i] = s[i] XOR cache[line a][i] for i in 0..15
|
||||
s = P_8(s)
|
||||
item(t) = s
|
||||
```
|
||||
|
||||
The 8 dependent cache reads per item are exactly today's: the address of read `r` is `s[0]` after program `r`, and
|
||||
`s[0]` depends on every earlier read through the chain (every round program writes every register, 2.4, and the
|
||||
XOR of the line into all 16 words feeds the next program). The verifier's latency part (8 dependent misses per
|
||||
item, overlapped across the up to 32 items of a load, spec 1.11) is unchanged, which the measurement shows: the
|
||||
difference against x8 is the ALU part only (section 5.1).
|
||||
|
||||
## 3. The prototype
|
||||
|
||||
### 3.1 Where it lives
|
||||
|
||||
| Item | Where |
|
||||
|---|---|
|
||||
| `DOp`, `DInstr`, `DeriveProgram` (draw, check, counts, fingerprint), the SoA interpreter `run_round` (pair dispatch), the scalar reference `run_round_scalar`, the text forms `instr_text` and `instr_line` | `igneum-pow/src/derive.rs` |
|
||||
| `Shape::derive_len`, `Shape::is_derived`, `MixParams::derive` (drawn after the 40 mixer draws), `derive_items` dispatching to `derive_items_program` | `igneum-pow/src/memhard.rs` |
|
||||
| `LoadClass::derive_len`, `LoadClass::DR736`, `with_derive`, parse and name `dr<len>`, the program id (`derive/` + the length) | `igneum-pow/src/generator.rs` |
|
||||
| `mh_round_0..8` and the program-driven `mh_item` in memhard.h, memhard.metal and kernel.cl; `IGNEUM_DERIVE_*` in program.h; `derive_len`, the op mix, the floors' counts and the nine programs (one line per instruction) in program.json | `igneum-pow/src/emit.rs` |
|
||||
| `--class dr736` (or any `dr<len>`) on every command; the bench prints the program's counts | `igneum-pow/src/main.rs` |
|
||||
| `examples/derive_perf.rs`: the interpreter's cost per instruction per batch, drawn program against uniform programs | `igneum-pow/examples/` |
|
||||
| Packs `dr736-genesis` (seed igneum-genesis, day 2026-10-03, program id 72c1d8048aef9542, program fingerprint 463535d01511350d) and `dr736-devnet-epoch0` (the devnet epoch 0 and day seeds, 7f4a5ca0a3637820, 771868df4e64d6ab); generator 2 with the class in the id, the x4-record shape, not a class v3 pack | `proto-cuda/packs-ca3-derive/` |
|
||||
| Tests: `src/derive.rs` (5), `tests/derive.rs` (7: by hand on a small cache, batches, v2 and v3 untouched, the stream and the class, determinism and the pack text, stats beside x8, the text forms against the scalar reference, the word path) | `igneum-pow` |
|
||||
| The PC 2 job | `relay/playbooks/ca3-derive-pc2.ps1` (section 5.4) |
|
||||
|
||||
### 3.2 The verifier without a JIT: the word-major interpreter
|
||||
|
||||
The verifier derives up to 32 distinct items per load (one per lane of the unit, `MemhardCpu::fetch`). The
|
||||
interpreter keeps the 32 item states word-major (`st[reg][lane]`, 2 KiB) and runs each instruction across the
|
||||
whole batch in one straight loop the compiler vectorises (NEON `add.4s`, `mul.4s`, `ushl.4s` and so on: 1,180
|
||||
such instructions in the example binary), so the dispatch is paid once per instruction per batch, not per item;
|
||||
the cache reads of the batch are issued together after each round program, as the fixed-mixer loop does, so the 8
|
||||
dependent misses of independent items overlap. The dispatch is on PAIRS of instructions (144 arms, one indirect
|
||||
branch per two instructions): a drawn op sequence is random, the predictor misses most dispatches, and pairing
|
||||
halves the misses per instruction. Measured with `examples/derive_perf.rs` (a functional run, load average 3.9 to
|
||||
4.9): 10.7 ns per instruction per batch cold, 7.18 warm with single dispatch, 4.98 with pair dispatch; a uniform
|
||||
program of one form (predictable dispatch) 3.4 to 4.3 ns, so the body is about 3.5 ns and the remaining dispatch
|
||||
cost about 1.5 ns. 6,624 x 4.98 ns x 128 batches = 4.2 ms per unit of interpreter time; the measured 4.88 ms
|
||||
includes the latency part and the transposes.
|
||||
|
||||
### 3.3 Bit-exactness on three vendors, by construction
|
||||
|
||||
Each form is one C statement on `uint` with `+`, `-`, `^`, `*`, `|`, `&` and the memhard core's `mh_rotl` (a
|
||||
shift pair, `n` in 1..31 at every call site), the same text in Metal, CUDA C and OpenCL C, every operand a 32-bit
|
||||
unsigned integer: the same argument as spec 1.14 for the lottery hash's families, which have run bit-exact on the
|
||||
three vendors since 4 October. The program is emitted as nine functions of straight-line statements (196 KB of
|
||||
memhard.h per day); NVRTC, the Metal compiler and the OpenCL compilers see no loop, no branch and no call inside a
|
||||
round program. Measured: section 5.2 (Metal and Apple OpenCL), 5.4 (CUDA).
|
||||
|
||||
### 3.4 The v2 and v3 paths are untouched
|
||||
|
||||
`Shape::derive_len` is 0 and `LoadClass::derive_len` is 0 on `V2`, `MX4`, `MX8` and `V3_CLASS`; `MixParams::derive`
|
||||
is `None`; `derive_items` takes the fixed-mixer loop as before; the emitter's text for a shape without a program is
|
||||
unchanged. `cargo test -p igneum-pow` (commit acb96ee): 58 lib, 7 derive, 4 mixer, 19 packs (every pinned pack of
|
||||
v2 and v3 regenerated and compared byte for byte), 7 scratch, all green.
|
||||
|
||||
## 4. What the chip keeps, and the fallbacks
|
||||
|
||||
### 4.1 The fixed-function allowance after the change
|
||||
|
||||
The chip of `chip-model-v3.md` now has to execute, per item, 6,624 instructions from a 12-form set over a
|
||||
16-entry register file, with three operand fields and a constant per instruction, in an order and with operands
|
||||
that change every day. That is a sequencer: an instruction store (6,624 x 9 bytes = 60 KB per day program, in
|
||||
SRAM beside the cache mirror), a register file with three read ports and one write port, a 32-bit ALU with a
|
||||
multiplier and a barrel rotator, operand muxes, and a program counter. A GPU streaming multiprocessor is the same
|
||||
machine with a wider register file and a warp scheduler. What the chip keeps over the GPU: no warp scheduler, no
|
||||
operand collector, no instruction cache hierarchy for a 1 MB kernel (the GPU's day program is about 60,000 SASS
|
||||
instructions per thread; whether the 5090's instruction cache holds it is in the build time of 5.4), no graphics
|
||||
or float units idle on the die, and the day's constants folded into the instruction store. What it loses: the
|
||||
wired pipeline (the mixer's 72 applications as 72 stages with the constants in the wires, no fetch, no register
|
||||
file, no crossbar), which is the thing the 3x credit paid for. The chain rule adds a second loss: with no
|
||||
intra-item parallelism a single engine finishes one instruction per cycle at best and its multiplies serialise
|
||||
unless it interleaves items, which costs a register file per item in flight (RandomX's light-mode argument,
|
||||
history 2.4: a chip paying "760 cycles and 1,240 multiplies per item"). ProgPoW claimed 1.1x to 1.2x for exactly
|
||||
this kind of chip ("conventional compute chips gain little on ProgPoW", Bob Rao's hardware audit, history 2.4,
|
||||
[S70]; EIP-1057's own claim 1.1x to 1.2x, [S67]). So the allowance this document carries is 1.2x (ProgPoW's
|
||||
claimed range, cited) with 1.5x as the cautious upper bound (approximate, mine), against the 3x of the fixed shape
|
||||
(approximate, from memory, M16). Section 7 prices all of 1.0x, 1.2x, 1.5x, 2x and 3x.
|
||||
|
||||
### 4.2 What it does not change
|
||||
|
||||
The partial-store chip (item 1): a chip that stores the dataset and never derives items pays nothing for the
|
||||
program; item 1's rows stand on their own. The cryptanalysis question (item 3) changes shape: instead of one
|
||||
fixed `M_r` to attack, the attacker gets a fresh random ARX program every day, which is RandomX's bet and is
|
||||
untested here; the day's programs can be audited by the same tools as random ARX ciphers, and the acceptance test
|
||||
is the place to add a structural rule if one is found. The era draws and the cache growth are untouched.
|
||||
|
||||
### 4.3 The JIT, named as the fallback with its risk
|
||||
|
||||
A per-day JIT for the verifier (emit NEON or AVX2 code for the 32-lane batch, the chain row kept in registers
|
||||
across instructions since `c` is always the row just written) would remove the 1.5 ns dispatch and about a third
|
||||
of the 3.5 ns body (the chain row's loads), about 2.5 ms per unit on this core, the x8 figure. Its risk: a code
|
||||
generator in the consensus path on two architectures (arm64, x86-64) whose output must equal the interpreter's
|
||||
bit for bit, writable-executable memory in a node and in every pool verifier, a new attack surface the history's
|
||||
lesson 9 says to audit before launch, and a second implementation per platform to keep in lockstep. Out of scope
|
||||
for this item; it is the route to 736 under the gate on a laptop core if the measurement of O-1.14 confirms the
|
||||
approximate row, and `dr368` is the route that needs no JIT.
|
||||
|
||||
## 5. Measurements
|
||||
|
||||
All under the locks of the brief; every row says its lock and the load average. Machine: Apple M5 Max, 64 GiB,
|
||||
Darwin 25.6.0. Commit acb96ee (the code and the packs), packbench built from this worktree.
|
||||
|
||||
### 5.1 The verifier per unit, one M5 Max core (`with-lock.sh measure`, one session, 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end)
|
||||
|
||||
Script `measure-ca3-derive.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03
|
||||
--class <v2|mx8|dr736> --warps 50`, two rounds, then the devnet seeds (`--epoch-hex edc4fa84...fb07 --day-hex
|
||||
69676e65756d2d6461792ffa50000000000000`) and `dr368` once.
|
||||
|
||||
| Class | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | Against x8 | Items per unit | Ops per item (GPU / chip / multiplies) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| v2 (igneum-genesis) | 0.598 / 0.594 | 0.697 | 1 | | 4,096 | 1,296 / 1,152 / 144 (9 x 144, hoisted 9 x 128) |
|
||||
| x8, mx8 (class v3) | 2.061 / 2.063 | 2.179 | 3.46x | 1 | 4,096 | 10,368 / 9,216 / 1,152 |
|
||||
| **dr736** | **4.875 / 4.944** | **5.241** | 8.2x | 2.37x | 4,096 | 10,659 / 9,992 / 1,461 |
|
||||
| x8, the devnet seeds | 2.078 | 2.155 | | | 4,095 to 4,096 | |
|
||||
| dr736, the devnet seeds | 4.872 | 5.241 | | 2.34x | 4,096 | 10,701 / 10,083 / 1,362 |
|
||||
| dr368 (the x4-equivalent fallback) | 2.692 | 2.898 | 4.5x | 1.31x | 4,096 | 5,350 / 5,004 / 752 |
|
||||
|
||||
Reading: the v2 row reads 0.59 to 0.60, the quiet readwidth night's 0.604 to 0.626 and 6.4a's 0.607 to 0.611, so
|
||||
this session is a quiet-core figure and no scaling applies. At the same chip-op count as x8 (plus 8%), the
|
||||
interpreter costs 2.4x what the compiled mixer costs: 4.98 ns per instruction per batch (3.2), of which about
|
||||
1.5 ns is dispatch and 3.5 ns the vector body, against the mixer's compiled straight-line loop over the same
|
||||
batch. The latency part is the same in every row (the 8 dependent misses per item; the dr368 row at half the
|
||||
instructions saves 2.2 ms of the 4.9, which puts the latency-and-transpose share at about 0.5 ms). The 10 ms gate
|
||||
keeps 5.1 ms (worst cold 4.76 ms) at 736 on this core and 7.3 ms at 368.
|
||||
|
||||
### 5.1a Verification throughput per tier (the form of mixer-x4.md 6.5)
|
||||
|
||||
| Figure | x8 (this session) | dr736 | dr368 | Note |
|
||||
|---|---|---|---|---|
|
||||
| ms per unit, quiet M5 Max core (measured, this session) | 2.06 | 4.88 | 2.69 | |
|
||||
| ms per unit, 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) | 5.2 | 12.2 | 6.7 | the figure that fixes the gate is a measurement, not this row |
|
||||
| Shares per second per core (quiet M5 Max) | 485 | 205 | 372 | |
|
||||
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second) | 4.5 | 10.7 | 5.9 | a pool verifying two units at once would halve the dispatch share (3.2); a 64-lane interpreter is a follow-up, unmeasured |
|
||||
| Node: worst cold single unit (per block) | 2.2 ms | 5.2 ms | 2.9 ms | a block's verification stays under the 1 s block time by 190x |
|
||||
| IBD over 108,000 headers on one core | 3.7 min | 8.8 min | 4.8 min | laptop (approximate): 9.4 / 22 / 12 min |
|
||||
| Margin left under the 10 ms gate (worst cold, this core) | 7.8 ms | 4.8 ms | 7.1 ms | on the laptop row (approximate): 4.8 / none (over by 2.2) / 3.3 ms |
|
||||
|
||||
Consequences per tier: every miner tier is untouched by the verifier (the miner never runs it); a pool operator
|
||||
pays 2.4x the cores at 736 (11 cores for a 22,000-member pool against 4.5) or 1.3x at 368; a node on any 2026 core
|
||||
verifies a block in 5 ms; a node on a 2019-class laptop core is the open question (O-1.14), and at 736 the
|
||||
approximate row says it misses the gate, so 736 does not go genesis-live on an approximation, and 368 passes it.
|
||||
|
||||
### 5.2 Bit-exactness on the Mac (`with-lock.sh run`, 08:38 to 08:41 local)
|
||||
|
||||
`packbench --pack <dir> --batches 1 --batch-log2 24 --group 256` (Metal, built from this worktree) and
|
||||
`igneum-bench-cl-dr736-genesis --bench-pack --pack <dir> --batches 1 --batch-log2 24` (Apple OpenCL, `proto-opencl/
|
||||
build.sh` on the pack through a temporary link). Vectors are the Rust interpreter's.
|
||||
|
||||
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word [MASK], 64 samples | Vectors | Fingerprint 2^24 | Compile | 1 GiB build, GPU ms (run lock, indicative) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| dr736-genesis | Metal | 48c4f5bf24166b2e PASS | head and last PASS (packbench checks no samples) | 3/3 standalone, 3/3 in batch | 50e3eaa779da4f1e | 784 ms (276 on the second process, 1 ms once the shader cache has it) | 38.2 / 38.3 |
|
||||
| dr736-genesis | Apple OpenCL | PASS (head, last line, FNV) | head, word [268435455], 64 samples PASS | 96 of 96 lanes | 50e3eaa779da4f1e | (in the 385 ms prepare) | 57 ms wall |
|
||||
| dr736-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 9553f6d5c667205a | 763 ms | 37.4 |
|
||||
|
||||
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the derived dataset (head, word [MASK], the 64
|
||||
samples through OpenCL), on every vector lane and on the 2^24-output fingerprint of dr736-genesis across both
|
||||
compilers; the devnet-seed pack agrees on Metal. The Apple OpenCL compile of a 6,624-statement item function
|
||||
is inside a 385 ms prepare; the Metal compile is 0.75 to 0.8 s cold per day (the memhard library is the day's,
|
||||
compiled once a day, not per epoch) and 1 ms from the shader cache.
|
||||
|
||||
### 5.3 The daily build and the hash rate on the M5 Max (`with-lock.sh measure`, the session of 5.1, three rounds)
|
||||
|
||||
`packbench --pack <dir> --batches 2 --batch-log2 22 --group 256`, mx8-genesis then dr736-genesis, three rounds.
|
||||
|
||||
| Pack | Compile (round 1 / 2 / 3) | 1 GiB build, GPU ms (round 1 / 2 / 3) | MH/s GPU (round 1 / 2 / 3) | Vectors, self-tests |
|
||||
|---|---|---|---|---|
|
||||
| mx8-genesis (class v3, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | 3/3 + 3/3, PASS |
|
||||
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | 3/3 + 3/3, PASS |
|
||||
|
||||
Reading: the build is 29 ms against 22 (+32%, +7 ms): the Mac's build was latency-bound at x1, x4 and x8 (21 to
|
||||
22 ms at every multiplier, mixer-x4.md 6.4) and the serial chain per thread now shows (no intra-item ILP for the
|
||||
compiler to schedule), still 34x under the 1 s bar. The hash rate is the x8 rate within 0.3% (27.06 to 27.13
|
||||
against 27.08 to 27.16), as it must be: the hash kernel only loads. Consequences per tier: a daily build of 29 ms
|
||||
on Apple silicon costs nothing to any tier; the 5090's figure is section 5.4, the 9070 XT's is OWED (PC 1 not
|
||||
released today; its x8 build was 72 to 77 ms and arithmetic-bound at none of x1, x4, x8, so the chain's cost there
|
||||
is the open number); the integrated tier is the one to watch (5.5).
|
||||
|
||||
### 5.4 RTX 5090, PC 2 (one job, `relay/playbooks/ca3-derive-pc2.ps1`)
|
||||
|
||||
PENDING at the time of writing: the job waits for `/tmp/igneum-devnet/pc2-ca3.clear` (the proving agent's
|
||||
30-minute measurement) and the `pc2-ca3.lock`. The job downloads the packs zip itself (sha256
|
||||
aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676: dr736-genesis, dr736-devnet-epoch0, mx8-genesis,
|
||||
v2-genesis-mh), runs the installed app's igneum-worker-cuda.exe (NVRTC compiles each pack's own text) with the
|
||||
NVIDIA card off in the app only under test (its key from settings.json), `--check` for the `nvrtc .. cache ..
|
||||
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
|
||||
The rows land in the bench-log entry when the closing report is read.
|
||||
|
||||
### 5.5 The build per tier, with the integrated tier
|
||||
|
||||
| Card | Build at x8 | Build with the day program | Source |
|
||||
|---|---|---|---|
|
||||
| M5 Max, Metal | 22.1 ms | 29.0 ms (+32%) | 5.3, measure lock |
|
||||
| RTX 5090, CUDA | 23 ms | section 5.4 | the PC 2 job |
|
||||
| RX 9070 XT, OpenCL | 72 to 77 ms | OWED (PC 1) | |
|
||||
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | about 55 to 94 s (approximate, mixer-x4.md 6.5: the iGPU's build is arithmetic-bound at x1 already, scaled x8 from 6.9 / 9.4 / 11.7 s) | about the same count of ops at a lower ILP: 55 to 120 s (approximate, unmeasured) | `docs/plans/epoch-length.md` 6.1 iGPU rows |
|
||||
| gfx1036 beside WSL build jobs (PC 1) | about 7 to 17 min (approximate) | the same or worse (approximate) | epoch-length.md 6.1 |
|
||||
| 8 GB-class discrete card (not owned, about a tenth of the 5090, approximate) | about 1 s | about 1 to 1.3 s (approximate) | scaled |
|
||||
|
||||
Consequences: nothing changes for a discrete card of any size on any vendor (the build is under a second), the
|
||||
Mac row measured; the integrated tier already misses the per-prepare rule at x8 and needs the per-day dataset
|
||||
reuse in the workers (0.3.12) or a restart per epoch, and the day program makes that need the same, not larger in
|
||||
kind; the per-day compile of the item function (0.75 s Metal; NVRTC on the 5090 in 5.4) lands once a day in the
|
||||
worker's day-cache build, not per epoch, unless the hash kernel's compile includes memhard.h, which on the CUDA
|
||||
worker it does (kernel_bound.cu includes it, and the variant race compiles 17 variants): the 5.4 job's nvrtc line
|
||||
is the number for that, and if it is large the fix is to compile the item function once per day into its own
|
||||
module.
|
||||
|
||||
## 6. PROPOSED spec text for 1.13.2: reserve entry R0, `derive` (the per-day item-derivation program)
|
||||
|
||||
Not written into `docs/spec`; it lives here until the project lead's word. Named R0, ahead of R1 (mm8), because the reserve
|
||||
is to be ordered by chip-unfriendliness (counter-asic-3.md item 6; mm8 last) and a derivation program is the most
|
||||
chip-unfriendly entry the reserve can hold: it removes the fixed-function allowance of the recompute chip rather
|
||||
than adding a family that chip can license.
|
||||
|
||||
> Reserve entry R0, `derive` (the per-day item-derivation program). Semantics: section 1.8.5 under
|
||||
> `derive_len = 736`: the nine mixer slots of the item derivation (one before each of the 8 cache reads, one after
|
||||
> the last) each run a straight-line program of 736 instructions drawn from the day key stream of 1.8.4 after its
|
||||
> 40 draws, four draws per instruction (`below(100)` the form, `below(15)` the destination among the registers
|
||||
> other than the chain, `below(14)` the third register among those other than the destination and the chain,
|
||||
> `next()` the immediate: `1 + low32 mod 31` for the rotate forms, `low32` for `addc` and `xorc`, `low32 OR 1` for
|
||||
> `mulc` and `mulc2`); the chain is `s[0]` at the start of each program and the previous destination after; the
|
||||
> twelve forms and weights of `docs/plans/counter-asic-3-derivation.md` section 2.2, fixed; every form a
|
||||
> bijection on the state; the 8 dependent cache reads, the mixer constants of the item init, the cache and the
|
||||
> dataset mapping of 1.8.5 unchanged; `mixer_mult` unused under R0. Acceptance test, per candidate, the next
|
||||
> attempt on rejection (the stream continues): every register written in every round program; at least 8 distinct
|
||||
> rotation amounts; per item at least 9,216 operations with constants folded, 10,368 as written and 1,152
|
||||
> multiplies (the x8 mixer's counts from `memhard::mixer`: 72 x 128, 72 x 144, 72 x 16; the floors scale with the
|
||||
> length). Edge vectors, each a hand-built item run on every vendor: item 0, item 1, item 2^28 - 1 and item 2^32 - 1
|
||||
> of the genesis day on a 2^16-word cache; a program whose first instruction is each of the twelve forms with
|
||||
> `d = 15`, `c = 0`, `b = 14`, `k = 31`, `i = 0xffffffff` (odd for the multiplies) on the all-ones state and on the
|
||||
> all-zero state (the wrap of every form); the day of the pinned pack `dr736-genesis` (program fingerprint
|
||||
> 463535d01511350d, dataset head `vectors.json`, 2^24 fingerprint 50e3eaa779da4f1e) and of `dr736-devnet-epoch0`
|
||||
> (771868df4e64d6ab, 9553f6d5c667205a). Unlock: at the start of era n = 2 (DAA 31,104,000), or earlier by the 90%
|
||||
> signalling path of section 5.7, or at genesis if the verifier on a 2019-class core (O-1.14) reads under 10 ms per
|
||||
> unit at 736, else at the length that does (368 measured at 2.69 ms on an M5 Max core); never by a release. The
|
||||
> verifier procedure: the word-major interpreter of `igneum-pow/src/derive.rs` (32 item states per batch, pair
|
||||
> dispatch), 4.88 ms per unit on one M5 Max core (section 5.1), no JIT; a JIT is the named fallback (4.3). Vendor
|
||||
> paths: none needed; every form is a single 32-bit integer statement on all three compilers (3.3).
|
||||
|
||||
## 7. The chip model row
|
||||
|
||||
Written into `docs/analysis/chip-model-v3.md` section 6 in the form of its section 2. Ops per hash: 128 items x
|
||||
9,992 chip ops (the genesis day's draw; the floor 9,216) = 1,278,976 (floor 1,179,648); chip rate at 50 T op/s =
|
||||
39.1 MH/s (floor 42.4); bare against 136.1 MH/s = 0.287x (floor 0.31x, the x8 row's figure, as the floor is x8's
|
||||
count). The allowance rows: 1.0x 0.29x; 1.2x (ProgPoW's claim) 0.34x; 1.5x (cautious upper bound, approximate)
|
||||
0.43x; 2x 0.57x; 3x (the fixed shape's, which no longer applies) 0.86x. Equal silicon (x 0.829): 0.24 / 0.29 /
|
||||
0.36 / 0.48 / 0.71. The x8 row read 0.92x at 3x and 0.76x at equal silicon; at the same 0.31x bare the day program
|
||||
takes the chip from 0.92x to 0.34x to 0.43x, which is the margin the item was for. dr368 (the fallback): 639,488
|
||||
chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x, with less margin than x8 had at
|
||||
3x and more than x4 had (1.84x).
|
||||
|
||||
## 8. What is unverified or owed
|
||||
|
||||
| Item | State |
|
||||
|---|---|
|
||||
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
|
||||
| RX 9070 XT (PC 1) | OWED: PC 1 is the project lead's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
|
||||
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
|
||||
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |
|
||||
| The integrated tier's build with the day program | approximate (5.5); the gfx1036 measurement is a PC 2 OpenCL job, not run today (the one PC 2 job carries the 5090) |
|
||||
| A 64-lane interpreter for pools (two units per batch) | unimplemented; it would cut the dispatch share for pool verifiers only |
|
||||
| The NVRTC cost of memhard.h inside the per-epoch hash kernel compile and the variant race | the PC 2 job's nvrtc line; the fix, if large, is one module per day for the item function |
|
||||
118
relay/playbooks/ca3-derive-pc2.ps1
Normal file
118
relay/playbooks/ca3-derive-pc2.ps1
Normal file
|
|
@ -0,0 +1,118 @@
|
|||
# Igneum run job: Counter ASIC 3.0 item 2, the per-day item-derivation program (docs/plans/counter-asic-3-derivation.md),
|
||||
# on PC 2's RTX 5090 (machine 1ccfe586), 6 October 2026. ONE job carries everything (the PC 2 rule of the ca3 brief):
|
||||
# it downloads the packs zip itself from the downloads host (sha256 checked), finds the installed app's
|
||||
# igneum-worker-cuda.exe (NVRTC compiles each pack's own kernel text, so no new worker build is needed), switches the
|
||||
# NVIDIA card off in the app ONLY while the packs run (the card key from the app's settings.json, never /api/state;
|
||||
# restored after with the settings it had), and never quits, restarts or updates the installed app. The numbers this
|
||||
# job is for: per pack, the worker's own `nvrtc .. cache .. dataset .. ms` line (the daily 1 GiB build on the 5090),
|
||||
# the self-test (bit-exactness: cache FNV, dataset head, word [MASK], 64 samples, 96 vector lanes against the Rust CPU
|
||||
# interpreter), the 2^24 fingerprint at base nonce 0 (against the Mac: dr736-genesis 50e3eaa779da4f1e,
|
||||
# dr736-devnet-epoch0 9553f6d5c667205a, mx8-genesis 7c28cfb06c5c65a9, v2-genesis-mh 25f96e7dce90bd4e) and the hash rate.
|
||||
# Packs: dr736-genesis, dr736-devnet-epoch0 (the derivation class), mx8-genesis (the x8 control), v2-genesis-mh (the v2
|
||||
# control). Every result line starts with RESULT. The placeholder __DL_BASE__ is substituted at publish time; the
|
||||
# downloads token never enters the repository.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
# everything lives in this job's own folder (the wiped-jobs-folder rule of tools/ci/kit-path-check.sh: nothing of
|
||||
# another job's is reached)
|
||||
$work = Join-Path $env:IGNEUM_JOB_DIR 'ca3-derive'
|
||||
New-Item -ItemType Directory -Force -Path $work | Out-Null
|
||||
$zipUrl = '__DL_BASE__/igneum-ca3-derive-packs.zip'
|
||||
$zipSha = 'aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676'
|
||||
$zip = Join-Path $work 'igneum-ca3-derive-packs.zip'
|
||||
Write-Output "RESULT job ca3-derive-pc2 start $(Stamp) machine $env:COMPUTERNAME"
|
||||
|
||||
# ---- the packs: downloaded by this job, sha256 checked, expanded fresh ----
|
||||
try {
|
||||
[Net.ServicePointManager]::SecurityProtocol = [Net.SecurityProtocolType]::Tls12
|
||||
Invoke-WebRequest -Uri $zipUrl -OutFile $zip -TimeoutSec 120 -Headers @{ 'Cache-Control' = 'no-cache' }
|
||||
} catch { Write-Output ("RESULT error download: " + $_.Exception.Message); exit 2 }
|
||||
$got = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower()
|
||||
if ($got -ne $zipSha) { Write-Output "RESULT error zip sha256 $got expected $zipSha"; exit 2 }
|
||||
Write-Output "RESULT zip sha256 $got size $((Get-Item $zip).Length) OK"
|
||||
$packs = Join-Path $work 'packs-ca3-derive'
|
||||
if (Test-Path $packs) { Remove-Item -Recurse -Force $packs }
|
||||
Expand-Archive -Path $zip -DestinationPath $work -Force
|
||||
$packList = @('v2-genesis-mh', 'mx8-genesis', 'dr736-genesis', 'dr736-devnet-epoch0')
|
||||
foreach ($pk in $packList) {
|
||||
$d = Join-Path $packs $pk
|
||||
if (-not (Test-Path (Join-Path $d 'memhard.h'))) { Write-Output "RESULT error pack $pk missing after extract"; exit 2 }
|
||||
Write-Output ("RESULT pack $pk memhard.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'memhard.h')).Hash.ToLower() + " program.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'program.h')).Hash.ToLower())
|
||||
}
|
||||
|
||||
# ---- the worker: the installed app's igneum-worker-cuda.exe, run in place (its NVRTC DLLs sit beside it) ----
|
||||
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
|
||||
if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe'; exit 2 }
|
||||
$cuda = Join-Path $inst 'igneum-worker-cuda.exe'
|
||||
Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) nvrtc_dlls $((Get-ChildItem $inst -Filter 'nvrtc*.dll').Count)"
|
||||
$help = (& $cuda --help 2>&1 | Out-String)
|
||||
$hasBench = $help -match '--bench'
|
||||
Write-Output "RESULT worker-cuda has --bench: $hasBench"
|
||||
if (-not $hasBench) { Write-Output 'RESULT note the installed worker has no --bench; --check gives the build time and the self-test, no fingerprint and no hash rate' }
|
||||
|
||||
# ---- the app: the NVIDIA card's key and settings from settings.json; switched off through POST api/cards only ----
|
||||
$appDir = $env:IGNEUM_APP_DIR
|
||||
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
|
||||
$urlFile = Join-Path $appDir 'app.url'
|
||||
$url = $null
|
||||
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
|
||||
$sj = Join-Path $appDir 'settings.json'
|
||||
$cardKey = $null; $cardPref = $null
|
||||
if (Test-Path $sj) {
|
||||
try {
|
||||
$settings = Get-Content -LiteralPath $sj -Raw | ConvertFrom-Json
|
||||
if ($settings.cards) {
|
||||
foreach ($p in $settings.cards.PSObject.Properties) { if ($p.Name -like 'nvidia:*') { $cardKey = $p.Name; $cardPref = $p.Value; break } }
|
||||
}
|
||||
} catch { Say ("settings.json: " + $_.Exception.Message) }
|
||||
}
|
||||
if ($cardKey) {
|
||||
Write-Output ("RESULT card " + $cardKey + " enabled=" + $cardPref.enabled + " identities=" + $cardPref.identities + " power_pct=" + $cardPref.power_pct + " (settings.json)")
|
||||
} else {
|
||||
Write-Output 'RESULT card none in settings.json (no nvidia:* entry); the app keeps mining on the card and the numbers carry that load'
|
||||
}
|
||||
$cardOff = $false
|
||||
if ($cardKey -and $url) {
|
||||
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
|
||||
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
|
||||
$body = @{ cards = @(@{ key = $cardKey; enabled = $false; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; $cardOff = $true; Say 'card off requested' } catch { Write-Output ("RESULT error api/cards off: " + $_.Exception.Message) }
|
||||
# wait for the worker process to go, up to 90 s (the process list, not /api/state)
|
||||
$t = 0
|
||||
while ($t -lt 90) {
|
||||
Start-Sleep -Seconds 5; $t += 5
|
||||
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
|
||||
if (-not $w) { break }
|
||||
}
|
||||
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
|
||||
Write-Output ("RESULT card-off " + $cardKey + " after " + $t + " s, worker processes left " + (($w | Measure-Object).Count))
|
||||
Start-Sleep -Seconds 5
|
||||
}
|
||||
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
|
||||
|
||||
# ---- the runs: --check for the build line and the self-test, --bench for the fingerprint and the rate ----
|
||||
foreach ($pk in $packList) {
|
||||
$d = Join-Path $packs $pk
|
||||
Write-Output "RESULT run $pk check start $(Stamp)"
|
||||
& $cuda --check --pack $d 2>&1 | ForEach-Object { "RESULT check $pk $_" }
|
||||
Write-Output "RESULT run $pk check exit $LASTEXITCODE"
|
||||
if ($hasBench) {
|
||||
Write-Output "RESULT run $pk bench start $(Stamp)"
|
||||
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT bench $pk $_" }
|
||||
Write-Output "RESULT run $pk bench exit $LASTEXITCODE"
|
||||
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT bench8 $pk $_" }
|
||||
}
|
||||
}
|
||||
& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
|
||||
|
||||
# ---- restore the card with the settings it had ----
|
||||
if ($cardOff) {
|
||||
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
|
||||
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
|
||||
$en = $true; if ($null -ne $cardPref.enabled) { $en = [bool]$cardPref.enabled }
|
||||
$body = @{ cards = @(@{ key = $cardKey; enabled = $en; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $cardKey + " enabled=" + $en) } catch { Write-Output ("RESULT error card restore " + $cardKey + ": " + $_.Exception.Message) }
|
||||
}
|
||||
Write-Output "RESULT job ca3-derive-pc2 end $(Stamp)"
|
||||
exit 0
|
||||
Loading…
Reference in a new issue