Merge branch 'ca3-derive' into ca3-coord
# Conflicts: # docs/analysis/chip-model-v3.md
This commit is contained in:
commit
1c08438a12
38 changed files with 58874 additions and 18 deletions
|
|
@ -332,3 +332,46 @@ measurements land.
|
|||
check that they are the right order.
|
||||
- The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000
|
||||
program ops are estimates; the program-length lever is a design item with its own measurements, not a result.
|
||||
|
||||
## 6. The per-day derivation (item 2)
|
||||
|
||||
6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything
|
||||
PROPOSED, a prototype behind load class `dr736`). The fixed-shape mixer of section 2's rows is replaced by nine
|
||||
straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms,
|
||||
the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item
|
||||
derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program
|
||||
of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a
|
||||
32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's
|
||||
constants in the wires. The counts are from the code (`memhard::mixer`: 144 ops per application as written, 128
|
||||
with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1;
|
||||
the day program's floor is those counts, so the bare row cannot fall below x8's.
|
||||
|
||||
| Row | Derivation | Chip ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) | Allowance 1.5x (cautious upper bound, approximate) | The old 3x (the fixed shape's; does not apply) | Equal silicon, SRAM deducted, at 1.2x / 1.5x |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) | fixed mixer, 72 x 128 | 1,179,648 | 42.4 MH/s | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
|
||||
| **dr736, the genesis day's draw** (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) | the day program, 9 x 736 instructions | 1,278,976 | 39.1 MH/s | 256 MiB | 128 / $46 | 0.29x | 0.34x | 0.43x | 0.86x | 0.29x / 0.36x |
|
||||
| dr736 at the floor (a day whose draw sits exactly on the acceptance floor) | the day program | 1,179,648 | 42.4 | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
|
||||
| dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) | the day program, 9 x 368 | 640,512 | 78.1 | 256 MiB | 128 / $46 | 0.57x | 0.69x | 0.86x | 1.72x | 0.57x / 0.71x |
|
||||
| dr736 at year 4 (cache 512 MiB) | the day program | 1,278,976 | 39.1 | 512 MiB | 255 / $111 | 0.29x | 0.34x | 0.43x | 0.86x | 0.23x / 0.28x |
|
||||
|
||||
Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2
|
||||
= 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357.
|
||||
The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or
|
||||
divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in;
|
||||
with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function
|
||||
is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector,
|
||||
and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional
|
||||
compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious
|
||||
upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule
|
||||
(every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no
|
||||
intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave
|
||||
items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument,
|
||||
history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's
|
||||
2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged;
|
||||
the 5090's build and compile are the PC 2 job, the 9070 XT's OWED.
|
||||
|
||||
What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their
|
||||
margins apply); no chip has been priced for its instruction store or its register files per item in flight; the
|
||||
random ARX programs have had no cryptanalysis (the item 3 brief should name them beside `M_r`); the 2019-class
|
||||
core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0
|
||||
carries that verdict.
|
||||
|
|
|
|||
|
|
@ -1961,3 +1961,84 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep
|
|||
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
|
||||
|
||||
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.
|
||||
|
||||
## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation
|
||||
|
||||
Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and
|
||||
the chip row in `docs/plans/counter-asic-3-derivation.md` and `docs/analysis/chip-model-v3.md` section 6.
|
||||
Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x
|
||||
to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash
|
||||
idea, `superscalar.cpp` read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op
|
||||
count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build
|
||||
and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average.
|
||||
|
||||
The construction (class `dr736`, `igneum-pow/src/derive.rs`): nine straight-line programs of 736 instructions per
|
||||
item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64
|
||||
stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by `c|1`, rotate-add, xor-rotate,
|
||||
add-constant, xor-constant, the `M_r` form `(d ^ c) * odd`, `d * odd + c`, `d ^= c & b`, `d += c | b`), every
|
||||
instruction reading the register the previous one wrote (the chain, `s[0]` first) and writing another, every form
|
||||
a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations,
|
||||
or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written,
|
||||
1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624
|
||||
instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major
|
||||
interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT.
|
||||
|
||||
**Verifier per 32-lane unit, one core (`with-lock.sh measure`, one session 07:42:20 to 07:42:33 UTC, load
|
||||
average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; `igneum-pow bench --seed igneum-genesis
|
||||
--day 2026-10-03 --class <c> --warps 50`, two rounds, then the devnet seeds once):**
|
||||
|
||||
| Class | ms per unit, avg of 50 (round 1 / 2) | Worst cold of three | Against x8 | Ops per item (GPU / chip / mul) |
|
||||
|---|---|---|---|---|
|
||||
| v2 | 0.598 / 0.594 | 0.697 | | 1,296 / 1,152 / 144 |
|
||||
| x8 (mx8, class v3) | 2.061 / 2.063 | 2.179 | 1 | 10,368 / 9,216 / 1,152 |
|
||||
| dr736 | 4.875 / 4.944 | 5.241 | 2.37x | 10,659 / 9,992 / 1,461 |
|
||||
| x8, devnet seeds | 2.078 | 2.155 | | |
|
||||
| dr736, devnet seeds | 4.872 | 5.241 | 2.34x | 10,701 / 10,083 / 1,362 |
|
||||
| dr368 (half length, the x4-equivalent fallback) | 2.692 | 2.898 | 1.31x | 5,350 / 5,004 / 752 |
|
||||
|
||||
The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core
|
||||
figures. The interpreter's cost split (`examples/derive_perf.rs`, a functional run under the run lock, load 3.9 to
|
||||
4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair
|
||||
dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the
|
||||
vector body (NEON, 1,180 `.4s` instructions in the binary).
|
||||
|
||||
**Bit-exactness (`with-lock.sh run`):** dr736-genesis on Metal (`packbench --batches 1 --batch-log2 24`) cache
|
||||
FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint
|
||||
2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (`igneum-bench-cl-dr736-genesis --bench-pack`) the
|
||||
self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0
|
||||
on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset
|
||||
and on 2^24 outputs.
|
||||
|
||||
**Daily 1 GiB build and hash rate, Metal (`with-lock.sh measure`, the same session, `packbench --batches 2
|
||||
--batch-log2 22 --group 256`, three rounds):**
|
||||
|
||||
| Pack | Compile (1 / 2 / 3) | Build, GPU ms (1 / 2 / 3) | MH/s GPU (1 / 2 / 3) |
|
||||
|---|---|---|---|
|
||||
| mx8-genesis (x8, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 |
|
||||
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 |
|
||||
|
||||
**Chip model (chip-model-v3.md section 6):** 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at
|
||||
50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no
|
||||
longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a
|
||||
cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.
|
||||
|
||||
**Consequences per tier.** The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms
|
||||
(x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs
|
||||
11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core
|
||||
(2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The
|
||||
build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below;
|
||||
the 9070 XT is OWED (PC 1 is the project lead's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier
|
||||
already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to
|
||||
94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the
|
||||
same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache;
|
||||
the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line
|
||||
is the number to read. The hash rate: unchanged within 0.3% on the Mac, as the hash kernel only loads. Packs grow
|
||||
by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.
|
||||
|
||||
**Go / no-go:** GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in
|
||||
docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms;
|
||||
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
|
||||
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
|
||||
|
||||
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
|
||||
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.
|
||||
|
|
|
|||
415
docs/plans/counter-asic-3-derivation.md
Normal file
415
docs/plans/counter-asic-3-derivation.md
Normal file
|
|
@ -0,0 +1,415 @@
|
|||
# Counter ASIC 3.0 item 2: a random item-derivation program per day
|
||||
|
||||
6 October 2026. Worker `derive` (branch `ca3-derive`), under the brief of `docs/plans/counter-asic-3.md` item 2 and
|
||||
the history audit's addition 2 (`docs/analysis/asic-resistance-history.md` section 4.3). Everything here is
|
||||
PROPOSED: a prototype behind a load class (`dr736`), measured on the Mac and on PC 2, written as a reserve entry for
|
||||
spec 1.13.2 (section 6) that lives in this file until the project lead's word. Nothing is published and no vector of class v2
|
||||
or v3 moves (section 3.4).
|
||||
|
||||
What it does, in one line: the fixed-shape mixer `M_r` of spec 1.8.4, applied 72 times per item under class v3,
|
||||
is replaced by nine straight-line programs of 736 instructions drawn once a day from the day key stream, so the chip
|
||||
that holds the cache on die must run an arbitrary program instead of a wired pipeline; the 8 dependent cache reads
|
||||
per item, the cache, the loads and the hash kernel are untouched.
|
||||
|
||||
## 0. The verdict first
|
||||
|
||||
| Gate | Number | Bar | Result |
|
||||
|---|---|---|---|
|
||||
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
|
||||
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
|
||||
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
|
||||
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is the project lead's desk today) |
|
||||
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
|
||||
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
|
||||
|
||||
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
|
||||
core measurement (O-1.14) lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core
|
||||
(pass) against about 12 ms on the approximate laptop row (fail). The class that passes both rows today is `dr368`
|
||||
(2.69 ms), at the x4-equivalent op count, with the chip row at 0.58x bare. The way to the 736 figure under the gate
|
||||
on a laptop is the JIT (section 4.3), which is out of scope and named with its risk.
|
||||
|
||||
## 1. Why: what the chip model says the fixed shape is worth
|
||||
|
||||
`docs/analysis/chip-model-v3.md` section 2 prices the on-die-cache recompute chip at 50 T op/s: class v3 (x8) costs
|
||||
it 1,198,080 integer ops per hash (72 mixer applications x 128 items x about 130 ops), 41.7 MH/s, 0.31x the 5090's
|
||||
136.1 MH/s bare, and 0.92x with "the 3x fixed-function factor (approximate, from memory: 2x to 5x is the usual
|
||||
credit for a pipeline with no scheduling or divergence)". That credit is the mixer's fixed shape: the chip unrolls
|
||||
the 72 applications into a wired pipeline with the day's constants baked in, no instruction fetch, no register
|
||||
file, no operand muxes. RandomX's answer (`vendor/RandomX/src/superscalar.cpp` is not in this tree; read on
|
||||
6 October 2026 from the upstream repository at commit 7607fb2 into the session scratchpad; the design argument is
|
||||
`doc/design.md`, history section 2.4) is SuperscalarHash: the dataset item derivation is itself a random program
|
||||
drawn from the cache key, 8 programs of about 450 instructions per item, so a light-mode chip "becomes a CPU".
|
||||
RandomX's generator (`generateSuperscalar`, lines 653 to 850) schedules for a superscalar x86 core: it picks a
|
||||
decode-buffer configuration per cycle, selects a source register that is ready at the cycle and a destination
|
||||
that is not the source and was not last written by the same op group (`selectDestination`, line 495: no
|
||||
"xor r,r2; xor r,r2", no "ror r,C1; ror r,C2", no two multiplies in a row on one register), and then computes the
|
||||
program's ASIC latency as the longest dependency chain (lines 810 to 824) and sets the address register to the
|
||||
register with the highest one. The item init is `rl[0] = (itemNumber + 1) * superscalarMul0`, the other seven
|
||||
registers `rl[0] ^ superscalarAdd_i`, then per cache access: run the program, XOR the mix block in, next address
|
||||
from the address register (`dataset.cpp` `initDatasetItem`, lines 164 to 190).
|
||||
|
||||
What carries over here and what does not: the per-day program, the acceptance by construction, the dependency
|
||||
chain and the address register idea carry over; the x86 port scheduling does not (our verifier interprets the
|
||||
program for 32 items at once and our miners compile it for a GPU, so the schedule that matters is the GPU's), and
|
||||
the latency bound RandomX relies on (the program's critical path against DRAM) is replaced by ours: the 8
|
||||
dependent cache reads per item, which are untouched.
|
||||
|
||||
## 2. The design
|
||||
|
||||
### 2.1 The draw
|
||||
|
||||
One SplitMix64 stream seeded with `K[0] | (K[1] << 32)` (the day key, spec 1.8.1), the 40 draws of 1.8.4
|
||||
(`ROT`, `MUL`, `RC`) first, exactly as today (the item init `s[8 + i] = t * MUL[i] + RC[i]` still uses them), then
|
||||
the program: `DERIVE_PROGRAMS = 9` round programs (one before each of the 8 cache reads, one after the last) of
|
||||
`derive_len` instructions each, four draws per instruction in a fixed order:
|
||||
|
||||
| Draw | Range | Sets |
|
||||
|---|---|---|
|
||||
| `below(100)` | op roll | the form, by the weight table of 2.2 (cumulative) |
|
||||
| `below(15)` | destination roll | `d`: the 15 registers other than the chain `c`, in ascending order (roll >= c adds one) |
|
||||
| `below(14)` | third-register roll | `b`: the 14 registers other than `d` and `c`, ascending (two skips); used by `andx` and `orx`, consumed by every form |
|
||||
| `next()` | the immediate | `k = 1 + (low32 mod 31)` for the rotate forms; `imm = low32` for `addc`, `xorc`; `imm = low32 OR 1` for `mulc`, `mulc2`; consumed by every form |
|
||||
|
||||
The chain register `c` is `s[0]` (the address word) at the start of each round program and the destination of the
|
||||
previous instruction after it. Every instruction reads `c` and writes `d != c`; the next instruction's chain is
|
||||
`d`. So no two instructions of a program can run in parallel (the "every instruction consumes the newest result"
|
||||
rule of the brief, SuperscalarHash's chain made strict), and no two consecutive instructions write one register,
|
||||
which is what removes the mergeable pairs SuperscalarHash's `selectDestination` guards against. Four draws per
|
||||
instruction whether the form uses them or not, so the stream position of every draw is fixed by its index and a
|
||||
future change to one form's draw changes no other draw (the rule 1.13.1 follows for `epoch_len`). A rejected
|
||||
candidate (2.4) is followed by the next: the stream continues, `attempt + 1`, as the program generator of 1.4.6.
|
||||
|
||||
Source: `igneum-pow/src/derive.rs` (`DeriveProgram::draw_candidate`, `draw`), `memhard.rs` (`MixParams::with_shape`).
|
||||
|
||||
### 2.2 The instruction set: twelve two-register forms, fixed at genesis
|
||||
|
||||
`c` the chain, `d` the destination, `b` the third register, `k` in 1..31, `i` a 32-bit constant (odd for the
|
||||
multiplies). All arithmetic modulo 2^32, no division, no float, no data-dependent branch. Every form is a bijection
|
||||
on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply, or a rotation of itself; `c`
|
||||
and `b` are not written), so a program loses no entropy, the property `M_r` has.
|
||||
|
||||
| Form | Semantics | Weight (percent) | GPU ops | Chip ops | Multiply |
|
||||
|---|---|---|---|---|---|
|
||||
| `add` | `d += c` | 14 | 1 | 1 | |
|
||||
| `sub` | `d -= c` | 10 | 1 | 1 | |
|
||||
| `xor` | `d ^= c` | 14 | 1 | 1 | |
|
||||
| `mul` | `d *= (c OR 1)` | 10 | 2 | 1 (the OR is a wire) | yes |
|
||||
| `rot` | `d = rotl(d, k) + c` | 10 | 2 | 2 | |
|
||||
| `xrot` | `d = rotl(d ^ c, k)` | 10 | 2 | 2 | |
|
||||
| `addc` | `d += c + i` | 6 | 2 | 2 | |
|
||||
| `xorc` | `d ^= c ^ i` | 6 | 2 | 2 | |
|
||||
| `mulc` | `d = (d ^ c) * i` (the per-word form of `M_r`, the chain in place of the round constant) | 8 | 2 | 2 | yes |
|
||||
| `mulc2` | `d = d * i + c` | 4 | 2 | 2 | yes |
|
||||
| `andx` | `d ^= (c AND b)` | 4 | 2 | 2 | |
|
||||
| `orx` | `d += (c OR b)` | 4 | 2 | 2 | |
|
||||
|
||||
Mean 1.62 GPU ops and 1.52 chip ops per instruction, 22% multiplies. `and` and `or` enter only as `andx` and
|
||||
`orx` (a destructive `d &= c` would lose bits; the XOR and add of a conjunction keep `d` invertible). The forms are
|
||||
the brief's set (add, sub, mul-lo, xor, rotate by 1..31, and, or, the ARX-multiply forms of `M_r`); `mulhi` is in
|
||||
the lottery hash's families and bit-exact on the three vendors, and is left out of the derivation on purpose so
|
||||
every form is one that the three compilers lower to a single integer instruction (section 3.3).
|
||||
|
||||
### 2.3 Length and the floors: the x8-equivalent op count
|
||||
|
||||
The coordinator's rule for item 2: the total op count per item equals or exceeds today's x8 count, so the chip's
|
||||
budget row does not fall. The x8 mixer counted from the code (`memhard::mixer`): 16 x (xor, add, mul) + 8 quarter
|
||||
rounds x 12 = 144 ops as written, 128 with the `RC[i] + rk` adds hoisted as constants (the chip and every compiler
|
||||
do that), 16 multiplies; 72 applications per item = 10,368 ops as written, 9,216 hoisted, 1,152 multiplies
|
||||
(`chip-model-v3.md` section 1 prices 130 per application from the spec text; the item 3 worker counted the same
|
||||
144 / 128, coordinator's note of 6 October). The floors are those three. `derive_len = 736` gives 6,624
|
||||
instructions per item, expected 10,731 GPU ops, 10,068 chip ops and 1,457 multiplies (the genesis day draws
|
||||
10,659 / 9,992 / 1,461; the devnet day 10,701 / 10,083 / 1,362), 7 to 22 standard deviations above the floors, so a
|
||||
rejection on a floor is a rare event and the test exists for the degenerate class. The class name is `dr736`; the
|
||||
half-length `dr368` (the x4-equivalent) is measured beside it as the fallback with its floors scaled.
|
||||
|
||||
### 2.4 The acceptance test
|
||||
|
||||
`DeriveProgram::check`, on every candidate; a rejection draws the next attempt from the stream:
|
||||
|
||||
| Test | Rejects | Expected rate at 736 |
|
||||
|---|---|---|
|
||||
| every register written in every round program | a register no instruction of a round program writes (its init word would never enter the chain within that round) | 16 x 9 x (14/15)^736 = under 10^-20 |
|
||||
| at least 8 distinct rotation amounts across the item's programs | all rotations equal (the brief's degenerate draw; the mixer's "ROT draw of eight equal values is possible and untested", spec 1.8.4) | about 1,300 rotate forms over 31 values: never |
|
||||
| chip ops >= 9,216, GPU ops >= 10,368, multiplies >= 1,152 per item (scaled to the length) | an op draw under the x8 count | 22, 9 and 9 standard deviations below the mean: never |
|
||||
| structural (asserted, hold by construction): `d != c`, `b` distinct from both, `k` in 1..31, odd multiplier constants | a generator bug | n/a |
|
||||
|
||||
The generator panics after 64 rejected candidates in a row (`MAX_ATTEMPTS`), which the rates above put beyond
|
||||
any day the chain will see; the panic is the right failure (a node that cannot derive the day's program cannot
|
||||
verify, and must say so rather than guess). The weights are fixed at genesis; only the order, the registers and
|
||||
the constants are drawn, so the family mix of a program cannot be steered by the draw.
|
||||
|
||||
### 2.5 The item, with the program in place
|
||||
|
||||
Spec 1.8.5 under the derivation class (`derive_len` nonzero, `mixer_mult` unused):
|
||||
|
||||
```
|
||||
s[i] = K[i] for i in 0..7
|
||||
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
|
||||
for r in 0..7:
|
||||
s = P_r(s) round program r, the chain starting at s[0]
|
||||
a = s[0] AND (2^(C - 4) - 1) cache line index, as today
|
||||
s[i] = s[i] XOR cache[line a][i] for i in 0..15
|
||||
s = P_8(s)
|
||||
item(t) = s
|
||||
```
|
||||
|
||||
The 8 dependent cache reads per item are exactly today's: the address of read `r` is `s[0]` after program `r`, and
|
||||
`s[0]` depends on every earlier read through the chain (every round program writes every register, 2.4, and the
|
||||
XOR of the line into all 16 words feeds the next program). The verifier's latency part (8 dependent misses per
|
||||
item, overlapped across the up to 32 items of a load, spec 1.11) is unchanged, which the measurement shows: the
|
||||
difference against x8 is the ALU part only (section 5.1).
|
||||
|
||||
## 3. The prototype
|
||||
|
||||
### 3.1 Where it lives
|
||||
|
||||
| Item | Where |
|
||||
|---|---|
|
||||
| `DOp`, `DInstr`, `DeriveProgram` (draw, check, counts, fingerprint), the SoA interpreter `run_round` (pair dispatch), the scalar reference `run_round_scalar`, the text forms `instr_text` and `instr_line` | `igneum-pow/src/derive.rs` |
|
||||
| `Shape::derive_len`, `Shape::is_derived`, `MixParams::derive` (drawn after the 40 mixer draws), `derive_items` dispatching to `derive_items_program` | `igneum-pow/src/memhard.rs` |
|
||||
| `LoadClass::derive_len`, `LoadClass::DR736`, `with_derive`, parse and name `dr<len>`, the program id (`derive/` + the length) | `igneum-pow/src/generator.rs` |
|
||||
| `mh_round_0..8` and the program-driven `mh_item` in memhard.h, memhard.metal and kernel.cl; `IGNEUM_DERIVE_*` in program.h; `derive_len`, the op mix, the floors' counts and the nine programs (one line per instruction) in program.json | `igneum-pow/src/emit.rs` |
|
||||
| `--class dr736` (or any `dr<len>`) on every command; the bench prints the program's counts | `igneum-pow/src/main.rs` |
|
||||
| `examples/derive_perf.rs`: the interpreter's cost per instruction per batch, drawn program against uniform programs | `igneum-pow/examples/` |
|
||||
| Packs `dr736-genesis` (seed igneum-genesis, day 2026-10-03, program id 72c1d8048aef9542, program fingerprint 463535d01511350d) and `dr736-devnet-epoch0` (the devnet epoch 0 and day seeds, 7f4a5ca0a3637820, 771868df4e64d6ab); generator 2 with the class in the id, the x4-record shape, not a class v3 pack | `proto-cuda/packs-ca3-derive/` |
|
||||
| Tests: `src/derive.rs` (5), `tests/derive.rs` (7: by hand on a small cache, batches, v2 and v3 untouched, the stream and the class, determinism and the pack text, stats beside x8, the text forms against the scalar reference, the word path) | `igneum-pow` |
|
||||
| The PC 2 job | `relay/playbooks/ca3-derive-pc2.ps1` (section 5.4) |
|
||||
|
||||
### 3.2 The verifier without a JIT: the word-major interpreter
|
||||
|
||||
The verifier derives up to 32 distinct items per load (one per lane of the unit, `MemhardCpu::fetch`). The
|
||||
interpreter keeps the 32 item states word-major (`st[reg][lane]`, 2 KiB) and runs each instruction across the
|
||||
whole batch in one straight loop the compiler vectorises (NEON `add.4s`, `mul.4s`, `ushl.4s` and so on: 1,180
|
||||
such instructions in the example binary), so the dispatch is paid once per instruction per batch, not per item;
|
||||
the cache reads of the batch are issued together after each round program, as the fixed-mixer loop does, so the 8
|
||||
dependent misses of independent items overlap. The dispatch is on PAIRS of instructions (144 arms, one indirect
|
||||
branch per two instructions): a drawn op sequence is random, the predictor misses most dispatches, and pairing
|
||||
halves the misses per instruction. Measured with `examples/derive_perf.rs` (a functional run, load average 3.9 to
|
||||
4.9): 10.7 ns per instruction per batch cold, 7.18 warm with single dispatch, 4.98 with pair dispatch; a uniform
|
||||
program of one form (predictable dispatch) 3.4 to 4.3 ns, so the body is about 3.5 ns and the remaining dispatch
|
||||
cost about 1.5 ns. 6,624 x 4.98 ns x 128 batches = 4.2 ms per unit of interpreter time; the measured 4.88 ms
|
||||
includes the latency part and the transposes.
|
||||
|
||||
### 3.3 Bit-exactness on three vendors, by construction
|
||||
|
||||
Each form is one C statement on `uint` with `+`, `-`, `^`, `*`, `|`, `&` and the memhard core's `mh_rotl` (a
|
||||
shift pair, `n` in 1..31 at every call site), the same text in Metal, CUDA C and OpenCL C, every operand a 32-bit
|
||||
unsigned integer: the same argument as spec 1.14 for the lottery hash's families, which have run bit-exact on the
|
||||
three vendors since 4 October. The program is emitted as nine functions of straight-line statements (196 KB of
|
||||
memhard.h per day); NVRTC, the Metal compiler and the OpenCL compilers see no loop, no branch and no call inside a
|
||||
round program. Measured: section 5.2 (Metal and Apple OpenCL), 5.4 (CUDA).
|
||||
|
||||
### 3.4 The v2 and v3 paths are untouched
|
||||
|
||||
`Shape::derive_len` is 0 and `LoadClass::derive_len` is 0 on `V2`, `MX4`, `MX8` and `V3_CLASS`; `MixParams::derive`
|
||||
is `None`; `derive_items` takes the fixed-mixer loop as before; the emitter's text for a shape without a program is
|
||||
unchanged. `cargo test -p igneum-pow` (commit acb96ee): 58 lib, 7 derive, 4 mixer, 19 packs (every pinned pack of
|
||||
v2 and v3 regenerated and compared byte for byte), 7 scratch, all green.
|
||||
|
||||
## 4. What the chip keeps, and the fallbacks
|
||||
|
||||
### 4.1 The fixed-function allowance after the change
|
||||
|
||||
The chip of `chip-model-v3.md` now has to execute, per item, 6,624 instructions from a 12-form set over a
|
||||
16-entry register file, with three operand fields and a constant per instruction, in an order and with operands
|
||||
that change every day. That is a sequencer: an instruction store (6,624 x 9 bytes = 60 KB per day program, in
|
||||
SRAM beside the cache mirror), a register file with three read ports and one write port, a 32-bit ALU with a
|
||||
multiplier and a barrel rotator, operand muxes, and a program counter. A GPU streaming multiprocessor is the same
|
||||
machine with a wider register file and a warp scheduler. What the chip keeps over the GPU: no warp scheduler, no
|
||||
operand collector, no instruction cache hierarchy for a 1 MB kernel (the GPU's day program is about 60,000 SASS
|
||||
instructions per thread; whether the 5090's instruction cache holds it is in the build time of 5.4), no graphics
|
||||
or float units idle on the die, and the day's constants folded into the instruction store. What it loses: the
|
||||
wired pipeline (the mixer's 72 applications as 72 stages with the constants in the wires, no fetch, no register
|
||||
file, no crossbar), which is the thing the 3x credit paid for. The chain rule adds a second loss: with no
|
||||
intra-item parallelism a single engine finishes one instruction per cycle at best and its multiplies serialise
|
||||
unless it interleaves items, which costs a register file per item in flight (RandomX's light-mode argument,
|
||||
history 2.4: a chip paying "760 cycles and 1,240 multiplies per item"). ProgPoW claimed 1.1x to 1.2x for exactly
|
||||
this kind of chip ("conventional compute chips gain little on ProgPoW", Bob Rao's hardware audit, history 2.4,
|
||||
[S70]; EIP-1057's own claim 1.1x to 1.2x, [S67]). So the allowance this document carries is 1.2x (ProgPoW's
|
||||
claimed range, cited) with 1.5x as the cautious upper bound (approximate, mine), against the 3x of the fixed shape
|
||||
(approximate, from memory, M16). Section 7 prices all of 1.0x, 1.2x, 1.5x, 2x and 3x.
|
||||
|
||||
### 4.2 What it does not change
|
||||
|
||||
The partial-store chip (item 1): a chip that stores the dataset and never derives items pays nothing for the
|
||||
program; item 1's rows stand on their own. The cryptanalysis question (item 3) changes shape: instead of one
|
||||
fixed `M_r` to attack, the attacker gets a fresh random ARX program every day, which is RandomX's bet and is
|
||||
untested here; the day's programs can be audited by the same tools as random ARX ciphers, and the acceptance test
|
||||
is the place to add a structural rule if one is found. The era draws and the cache growth are untouched.
|
||||
|
||||
### 4.3 The JIT, named as the fallback with its risk
|
||||
|
||||
A per-day JIT for the verifier (emit NEON or AVX2 code for the 32-lane batch, the chain row kept in registers
|
||||
across instructions since `c` is always the row just written) would remove the 1.5 ns dispatch and about a third
|
||||
of the 3.5 ns body (the chain row's loads), about 2.5 ms per unit on this core, the x8 figure. Its risk: a code
|
||||
generator in the consensus path on two architectures (arm64, x86-64) whose output must equal the interpreter's
|
||||
bit for bit, writable-executable memory in a node and in every pool verifier, a new attack surface the history's
|
||||
lesson 9 says to audit before launch, and a second implementation per platform to keep in lockstep. Out of scope
|
||||
for this item; it is the route to 736 under the gate on a laptop core if the measurement of O-1.14 confirms the
|
||||
approximate row, and `dr368` is the route that needs no JIT.
|
||||
|
||||
## 5. Measurements
|
||||
|
||||
All under the locks of the brief; every row says its lock and the load average. Machine: Apple M5 Max, 64 GiB,
|
||||
Darwin 25.6.0. Commit acb96ee (the code and the packs), packbench built from this worktree.
|
||||
|
||||
### 5.1 The verifier per unit, one M5 Max core (`with-lock.sh measure`, one session, 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end)
|
||||
|
||||
Script `measure-ca3-derive.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03
|
||||
--class <v2|mx8|dr736> --warps 50`, two rounds, then the devnet seeds (`--epoch-hex edc4fa84...fb07 --day-hex
|
||||
69676e65756d2d6461792ffa50000000000000`) and `dr368` once.
|
||||
|
||||
| Class | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | Against x8 | Items per unit | Ops per item (GPU / chip / multiplies) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| v2 (igneum-genesis) | 0.598 / 0.594 | 0.697 | 1 | | 4,096 | 1,296 / 1,152 / 144 (9 x 144, hoisted 9 x 128) |
|
||||
| x8, mx8 (class v3) | 2.061 / 2.063 | 2.179 | 3.46x | 1 | 4,096 | 10,368 / 9,216 / 1,152 |
|
||||
| **dr736** | **4.875 / 4.944** | **5.241** | 8.2x | 2.37x | 4,096 | 10,659 / 9,992 / 1,461 |
|
||||
| x8, the devnet seeds | 2.078 | 2.155 | | | 4,095 to 4,096 | |
|
||||
| dr736, the devnet seeds | 4.872 | 5.241 | | 2.34x | 4,096 | 10,701 / 10,083 / 1,362 |
|
||||
| dr368 (the x4-equivalent fallback) | 2.692 | 2.898 | 4.5x | 1.31x | 4,096 | 5,350 / 5,004 / 752 |
|
||||
|
||||
Reading: the v2 row reads 0.59 to 0.60, the quiet readwidth night's 0.604 to 0.626 and 6.4a's 0.607 to 0.611, so
|
||||
this session is a quiet-core figure and no scaling applies. At the same chip-op count as x8 (plus 8%), the
|
||||
interpreter costs 2.4x what the compiled mixer costs: 4.98 ns per instruction per batch (3.2), of which about
|
||||
1.5 ns is dispatch and 3.5 ns the vector body, against the mixer's compiled straight-line loop over the same
|
||||
batch. The latency part is the same in every row (the 8 dependent misses per item; the dr368 row at half the
|
||||
instructions saves 2.2 ms of the 4.9, which puts the latency-and-transpose share at about 0.5 ms). The 10 ms gate
|
||||
keeps 5.1 ms (worst cold 4.76 ms) at 736 on this core and 7.3 ms at 368.
|
||||
|
||||
### 5.1a Verification throughput per tier (the form of mixer-x4.md 6.5)
|
||||
|
||||
| Figure | x8 (this session) | dr736 | dr368 | Note |
|
||||
|---|---|---|---|---|
|
||||
| ms per unit, quiet M5 Max core (measured, this session) | 2.06 | 4.88 | 2.69 | |
|
||||
| ms per unit, 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) | 5.2 | 12.2 | 6.7 | the figure that fixes the gate is a measurement, not this row |
|
||||
| Shares per second per core (quiet M5 Max) | 485 | 205 | 372 | |
|
||||
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second) | 4.5 | 10.7 | 5.9 | a pool verifying two units at once would halve the dispatch share (3.2); a 64-lane interpreter is a follow-up, unmeasured |
|
||||
| Node: worst cold single unit (per block) | 2.2 ms | 5.2 ms | 2.9 ms | a block's verification stays under the 1 s block time by 190x |
|
||||
| IBD over 108,000 headers on one core | 3.7 min | 8.8 min | 4.8 min | laptop (approximate): 9.4 / 22 / 12 min |
|
||||
| Margin left under the 10 ms gate (worst cold, this core) | 7.8 ms | 4.8 ms | 7.1 ms | on the laptop row (approximate): 4.8 / none (over by 2.2) / 3.3 ms |
|
||||
|
||||
Consequences per tier: every miner tier is untouched by the verifier (the miner never runs it); a pool operator
|
||||
pays 2.4x the cores at 736 (11 cores for a 22,000-member pool against 4.5) or 1.3x at 368; a node on any 2026 core
|
||||
verifies a block in 5 ms; a node on a 2019-class laptop core is the open question (O-1.14), and at 736 the
|
||||
approximate row says it misses the gate, so 736 does not go genesis-live on an approximation, and 368 passes it.
|
||||
|
||||
### 5.2 Bit-exactness on the Mac (`with-lock.sh run`, 08:38 to 08:41 local)
|
||||
|
||||
`packbench --pack <dir> --batches 1 --batch-log2 24 --group 256` (Metal, built from this worktree) and
|
||||
`igneum-bench-cl-dr736-genesis --bench-pack --pack <dir> --batches 1 --batch-log2 24` (Apple OpenCL, `proto-opencl/
|
||||
build.sh` on the pack through a temporary link). Vectors are the Rust interpreter's.
|
||||
|
||||
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word [MASK], 64 samples | Vectors | Fingerprint 2^24 | Compile | 1 GiB build, GPU ms (run lock, indicative) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| dr736-genesis | Metal | 48c4f5bf24166b2e PASS | head and last PASS (packbench checks no samples) | 3/3 standalone, 3/3 in batch | 50e3eaa779da4f1e | 784 ms (276 on the second process, 1 ms once the shader cache has it) | 38.2 / 38.3 |
|
||||
| dr736-genesis | Apple OpenCL | PASS (head, last line, FNV) | head, word [268435455], 64 samples PASS | 96 of 96 lanes | 50e3eaa779da4f1e | (in the 385 ms prepare) | 57 ms wall |
|
||||
| dr736-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 9553f6d5c667205a | 763 ms | 37.4 |
|
||||
|
||||
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the derived dataset (head, word [MASK], the 64
|
||||
samples through OpenCL), on every vector lane and on the 2^24-output fingerprint of dr736-genesis across both
|
||||
compilers; the devnet-seed pack agrees on Metal. The Apple OpenCL compile of a 6,624-statement item function
|
||||
is inside a 385 ms prepare; the Metal compile is 0.75 to 0.8 s cold per day (the memhard library is the day's,
|
||||
compiled once a day, not per epoch) and 1 ms from the shader cache.
|
||||
|
||||
### 5.3 The daily build and the hash rate on the M5 Max (`with-lock.sh measure`, the session of 5.1, three rounds)
|
||||
|
||||
`packbench --pack <dir> --batches 2 --batch-log2 22 --group 256`, mx8-genesis then dr736-genesis, three rounds.
|
||||
|
||||
| Pack | Compile (round 1 / 2 / 3) | 1 GiB build, GPU ms (round 1 / 2 / 3) | MH/s GPU (round 1 / 2 / 3) | Vectors, self-tests |
|
||||
|---|---|---|---|---|
|
||||
| mx8-genesis (class v3, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | 3/3 + 3/3, PASS |
|
||||
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | 3/3 + 3/3, PASS |
|
||||
|
||||
Reading: the build is 29 ms against 22 (+32%, +7 ms): the Mac's build was latency-bound at x1, x4 and x8 (21 to
|
||||
22 ms at every multiplier, mixer-x4.md 6.4) and the serial chain per thread now shows (no intra-item ILP for the
|
||||
compiler to schedule), still 34x under the 1 s bar. The hash rate is the x8 rate within 0.3% (27.06 to 27.13
|
||||
against 27.08 to 27.16), as it must be: the hash kernel only loads. Consequences per tier: a daily build of 29 ms
|
||||
on Apple silicon costs nothing to any tier; the 5090's figure is section 5.4, the 9070 XT's is OWED (PC 1 not
|
||||
released today; its x8 build was 72 to 77 ms and arithmetic-bound at none of x1, x4, x8, so the chain's cost there
|
||||
is the open number); the integrated tier is the one to watch (5.5).
|
||||
|
||||
### 5.4 RTX 5090, PC 2 (one job, `relay/playbooks/ca3-derive-pc2.ps1`)
|
||||
|
||||
PENDING at the time of writing: the job waits for `/tmp/igneum-devnet/pc2-ca3.clear` (the proving agent's
|
||||
30-minute measurement) and the `pc2-ca3.lock`. The job downloads the packs zip itself (sha256
|
||||
aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676: dr736-genesis, dr736-devnet-epoch0, mx8-genesis,
|
||||
v2-genesis-mh), runs the installed app's igneum-worker-cuda.exe (NVRTC compiles each pack's own text) with the
|
||||
NVIDIA card off in the app only under test (its key from settings.json), `--check` for the `nvrtc .. cache ..
|
||||
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
|
||||
The rows land in the bench-log entry when the closing report is read.
|
||||
|
||||
### 5.5 The build per tier, with the integrated tier
|
||||
|
||||
| Card | Build at x8 | Build with the day program | Source |
|
||||
|---|---|---|---|
|
||||
| M5 Max, Metal | 22.1 ms | 29.0 ms (+32%) | 5.3, measure lock |
|
||||
| RTX 5090, CUDA | 23 ms | section 5.4 | the PC 2 job |
|
||||
| RX 9070 XT, OpenCL | 72 to 77 ms | OWED (PC 1) | |
|
||||
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | about 55 to 94 s (approximate, mixer-x4.md 6.5: the iGPU's build is arithmetic-bound at x1 already, scaled x8 from 6.9 / 9.4 / 11.7 s) | about the same count of ops at a lower ILP: 55 to 120 s (approximate, unmeasured) | `docs/plans/epoch-length.md` 6.1 iGPU rows |
|
||||
| gfx1036 beside WSL build jobs (PC 1) | about 7 to 17 min (approximate) | the same or worse (approximate) | epoch-length.md 6.1 |
|
||||
| 8 GB-class discrete card (not owned, about a tenth of the 5090, approximate) | about 1 s | about 1 to 1.3 s (approximate) | scaled |
|
||||
|
||||
Consequences: nothing changes for a discrete card of any size on any vendor (the build is under a second), the
|
||||
Mac row measured; the integrated tier already misses the per-prepare rule at x8 and needs the per-day dataset
|
||||
reuse in the workers (0.3.12) or a restart per epoch, and the day program makes that need the same, not larger in
|
||||
kind; the per-day compile of the item function (0.75 s Metal; NVRTC on the 5090 in 5.4) lands once a day in the
|
||||
worker's day-cache build, not per epoch, unless the hash kernel's compile includes memhard.h, which on the CUDA
|
||||
worker it does (kernel_bound.cu includes it, and the variant race compiles 17 variants): the 5.4 job's nvrtc line
|
||||
is the number for that, and if it is large the fix is to compile the item function once per day into its own
|
||||
module.
|
||||
|
||||
## 6. PROPOSED spec text for 1.13.2: reserve entry R0, `derive` (the per-day item-derivation program)
|
||||
|
||||
Not written into `docs/spec`; it lives here until the project lead's word. Named R0, ahead of R1 (mm8), because the reserve
|
||||
is to be ordered by chip-unfriendliness (counter-asic-3.md item 6; mm8 last) and a derivation program is the most
|
||||
chip-unfriendly entry the reserve can hold: it removes the fixed-function allowance of the recompute chip rather
|
||||
than adding a family that chip can license.
|
||||
|
||||
> Reserve entry R0, `derive` (the per-day item-derivation program). Semantics: section 1.8.5 under
|
||||
> `derive_len = 736`: the nine mixer slots of the item derivation (one before each of the 8 cache reads, one after
|
||||
> the last) each run a straight-line program of 736 instructions drawn from the day key stream of 1.8.4 after its
|
||||
> 40 draws, four draws per instruction (`below(100)` the form, `below(15)` the destination among the registers
|
||||
> other than the chain, `below(14)` the third register among those other than the destination and the chain,
|
||||
> `next()` the immediate: `1 + low32 mod 31` for the rotate forms, `low32` for `addc` and `xorc`, `low32 OR 1` for
|
||||
> `mulc` and `mulc2`); the chain is `s[0]` at the start of each program and the previous destination after; the
|
||||
> twelve forms and weights of `docs/plans/counter-asic-3-derivation.md` section 2.2, fixed; every form a
|
||||
> bijection on the state; the 8 dependent cache reads, the mixer constants of the item init, the cache and the
|
||||
> dataset mapping of 1.8.5 unchanged; `mixer_mult` unused under R0. Acceptance test, per candidate, the next
|
||||
> attempt on rejection (the stream continues): every register written in every round program; at least 8 distinct
|
||||
> rotation amounts; per item at least 9,216 operations with constants folded, 10,368 as written and 1,152
|
||||
> multiplies (the x8 mixer's counts from `memhard::mixer`: 72 x 128, 72 x 144, 72 x 16; the floors scale with the
|
||||
> length). Edge vectors, each a hand-built item run on every vendor: item 0, item 1, item 2^28 - 1 and item 2^32 - 1
|
||||
> of the genesis day on a 2^16-word cache; a program whose first instruction is each of the twelve forms with
|
||||
> `d = 15`, `c = 0`, `b = 14`, `k = 31`, `i = 0xffffffff` (odd for the multiplies) on the all-ones state and on the
|
||||
> all-zero state (the wrap of every form); the day of the pinned pack `dr736-genesis` (program fingerprint
|
||||
> 463535d01511350d, dataset head `vectors.json`, 2^24 fingerprint 50e3eaa779da4f1e) and of `dr736-devnet-epoch0`
|
||||
> (771868df4e64d6ab, 9553f6d5c667205a). Unlock: at the start of era n = 2 (DAA 31,104,000), or earlier by the 90%
|
||||
> signalling path of section 5.7, or at genesis if the verifier on a 2019-class core (O-1.14) reads under 10 ms per
|
||||
> unit at 736, else at the length that does (368 measured at 2.69 ms on an M5 Max core); never by a release. The
|
||||
> verifier procedure: the word-major interpreter of `igneum-pow/src/derive.rs` (32 item states per batch, pair
|
||||
> dispatch), 4.88 ms per unit on one M5 Max core (section 5.1), no JIT; a JIT is the named fallback (4.3). Vendor
|
||||
> paths: none needed; every form is a single 32-bit integer statement on all three compilers (3.3).
|
||||
|
||||
## 7. The chip model row
|
||||
|
||||
Written into `docs/analysis/chip-model-v3.md` section 6 in the form of its section 2. Ops per hash: 128 items x
|
||||
9,992 chip ops (the genesis day's draw; the floor 9,216) = 1,278,976 (floor 1,179,648); chip rate at 50 T op/s =
|
||||
39.1 MH/s (floor 42.4); bare against 136.1 MH/s = 0.287x (floor 0.31x, the x8 row's figure, as the floor is x8's
|
||||
count). The allowance rows: 1.0x 0.29x; 1.2x (ProgPoW's claim) 0.34x; 1.5x (cautious upper bound, approximate)
|
||||
0.43x; 2x 0.57x; 3x (the fixed shape's, which no longer applies) 0.86x. Equal silicon (x 0.829): 0.24 / 0.29 /
|
||||
0.36 / 0.48 / 0.71. The x8 row read 0.92x at 3x and 0.76x at equal silicon; at the same 0.31x bare the day program
|
||||
takes the chip from 0.92x to 0.34x to 0.43x, which is the margin the item was for. dr368 (the fallback): 639,488
|
||||
chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x, with less margin than x8 had at
|
||||
3x and more than x4 had (1.84x).
|
||||
|
||||
## 8. What is unverified or owed
|
||||
|
||||
| Item | State |
|
||||
|---|---|
|
||||
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
|
||||
| RX 9070 XT (PC 1) | OWED: PC 1 is the project lead's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
|
||||
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
|
||||
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |
|
||||
| The integrated tier's build with the day program | approximate (5.5); the gfx1036 measurement is a PC 2 OpenCL job, not run today (the one PC 2 job carries the 5090) |
|
||||
| A 64-lane interpreter for pools (two units per batch) | unimplemented; it would cut the dispatch share for pool verifiers only |
|
||||
| The NVRTC cost of memhard.h inside the per-epoch hash kernel compile and the variant race | the PC 2 job's nvrtc line; the fix, if large, is one module per day for the item function |
|
||||
51
igneum-pow/examples/derive_perf.rs
Normal file
51
igneum-pow/examples/derive_perf.rs
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
//! The derivation interpreter's cost per instruction per batch (Counter ASIC 3.0 item 2): the day's program against
|
||||
//! a uniform program of the same length (every instruction one form, so the dispatch is predictable), which splits
|
||||
//! the per-instruction cost into the dispatch and the vector body. A functional tool, not a bench-log number on its
|
||||
//! own: run it under the measure lock and state the load average when a figure is recorded.
|
||||
//! cargo run --release --example derive_perf [len] [reps]
|
||||
use igneum_pow::derive::{run_round, DInstr, DOp, DeriveProgram, SoaState, DERIVE_LEN_X8, DERIVE_REGS, SOA_LANES};
|
||||
use igneum_pow::seed::SplitMix64;
|
||||
use std::time::Instant;
|
||||
|
||||
fn time(prog: &DeriveProgram, reps: usize) -> (f64, u64) {
|
||||
let mut st: SoaState = [[0u32; SOA_LANES]; DERIVE_REGS];
|
||||
let mut x = SplitMix64::new(5);
|
||||
for r in 0..DERIVE_REGS {
|
||||
for k in 0..SOA_LANES {
|
||||
st[r][k] = x.next() as u32;
|
||||
}
|
||||
}
|
||||
let t = Instant::now();
|
||||
for _ in 0..reps {
|
||||
for p in &prog.rounds {
|
||||
run_round(p, &mut st);
|
||||
}
|
||||
}
|
||||
let ns = t.elapsed().as_nanos() as f64 / (reps as f64 * prog.instr_count() as f64);
|
||||
(ns, st[0][0] as u64 ^ st[15][31] as u64)
|
||||
}
|
||||
|
||||
fn uniform(len: u32, op: DOp) -> DeriveProgram {
|
||||
let mut p = DeriveProgram::draw_candidate(&mut SplitMix64::new(1), len, 0);
|
||||
for ins in p.rounds.iter_mut().flatten() {
|
||||
*ins = DInstr { op, rot: if op.has_rot() { 13 } else { 0 }, imm: if op.has_imm() { 0x9e37_79b9 } else { 0 }, ..*ins };
|
||||
}
|
||||
p
|
||||
}
|
||||
|
||||
fn main() {
|
||||
let a: Vec<String> = std::env::args().collect();
|
||||
let len: u32 = a.get(1).and_then(|s| s.parse().ok()).unwrap_or(DERIVE_LEN_X8);
|
||||
let reps: usize = a.get(2).and_then(|s| s.parse().ok()).unwrap_or(400);
|
||||
let real = DeriveProgram::draw(&mut SplitMix64::new(0x3067619f3c269176), len);
|
||||
let per_item = real.instr_count();
|
||||
println!("len {len}: {per_item} instructions per item, {} GPU ops, {} chip ops, {} multiplies; batch of {SOA_LANES} lanes, {reps} reps", real.gpu_ops(), real.chip_ops(), real.muls());
|
||||
for round in 0..2 {
|
||||
let (ns, sink) = time(&real, reps);
|
||||
println!("round {round}: drawn program {ns:.2} ns per instruction per batch ({:.1} us per batch, {:.2} ms per 128 batches) sink {sink:x}", ns * per_item as f64 / 1e3, ns * per_item as f64 * 128.0 / 1e6);
|
||||
for op in [DOp::Add, DOp::Mul, DOp::XRot, DOp::MulC, DOp::AndX] {
|
||||
let (ns, sink) = time(&uniform(len, op), reps);
|
||||
println!("round {round}: uniform {:5} {ns:.2} ns per instruction per batch sink {sink:x}", op.name());
|
||||
}
|
||||
}
|
||||
}
|
||||
978
igneum-pow/src/derive.rs
Normal file
978
igneum-pow/src/derive.rs
Normal file
|
|
@ -0,0 +1,978 @@
|
|||
//! The per-day item-derivation program (Counter ASIC 3.0 item 2, `docs/plans/counter-asic-3-derivation.md`):
|
||||
//! RandomX's SuperscalarHash idea (`vendor/RandomX/src/superscalar.cpp`, read 6 October 2026 at commit 7607fb2)
|
||||
//! rebuilt for a 16-word item on a GPU. In place of the fixed-shape mixer `M_r` of spec 01 section 1.8.4, each of
|
||||
//! the nine mixer slots of an item (one before each of the 8 dependent cache reads, one after the last) runs a
|
||||
//! straight-line program of [`DERIVE_LEN`] instructions drawn once a day from the day key stream, from a fixed
|
||||
//! set of twelve two-register forms. The 8 dependent cache reads per item are untouched.
|
||||
//!
|
||||
//! Rules of the draw (the dependency chain of SuperscalarHash, made strict):
|
||||
//! * every instruction reads the chain register `c`, the register the previous instruction wrote (`s[0]`, the
|
||||
//! address word, at the start of each round program), and writes a register `d != c`, which becomes the chain;
|
||||
//! so no two instructions of a program can run in parallel, and no two consecutive instructions write one
|
||||
//! register (the "ror r,C1; ror r,C2" and "xor r,r2; xor r,r2" merges of SuperscalarHash's `selectDestination`
|
||||
//! cannot arise);
|
||||
//! * every form is a bijection on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply,
|
||||
//! or a rotation of itself), so a program loses no entropy, the property `M_r` has;
|
||||
//! * the forms are integer only, modulo 2^32, with rotations by 1..31: bit-exact on Metal, CUDA and OpenCL by the
|
||||
//! same argument as the lottery hash's families (spec 01 section 1.14); no division, no float, no branch;
|
||||
//! * four draws per instruction in a fixed order, so the stream position of every draw is fixed by the index.
|
||||
//!
|
||||
//! The acceptance test ([`DeriveProgram::check`]) rejects a degenerate draw and the next attempt is drawn from the
|
||||
//! continuation of the stream, the rule the program generator uses (spec 01 section 1.4.6).
|
||||
|
||||
use crate::memhard::ITEM_ROUNDS;
|
||||
use crate::seed::SplitMix64;
|
||||
|
||||
/// Registers of the item state (the item is 16 words).
|
||||
pub const DERIVE_REGS: usize = 16;
|
||||
/// Round programs per item: one before each cache read and one after the last (`ITEM_ROUNDS + 1`).
|
||||
pub const DERIVE_PROGRAMS: usize = ITEM_ROUNDS + 1;
|
||||
/// Instructions per round program for the x8-equivalent operation count (the candidate, class "dr736"): 9 x 736
|
||||
/// = 6,624 instructions per item at a mean of 1.62 GPU operations (1.52 chip operations) each, about 10,730 GPU
|
||||
/// operations, 10,070 chip operations and 1,460 multiplies per item. The x8 mixer, counted from the code
|
||||
/// (`memhard::mixer`, 16 x (xor, add, mul) + 8 quarter rounds x 12 = 144 operations as written, 128 with the
|
||||
/// `RC + rk` adds hoisted as constants, 16 multiplies; `chip-model-v3.md` section 1 prices 130 from the spec text):
|
||||
/// 72 applications = 10,368 as written, 9,216 hoisted, 1,152 multiplies. The floors below are those three.
|
||||
pub const DERIVE_LEN_X8: u32 = 736;
|
||||
/// Draws per instruction: the op roll, the destination roll, the second-source roll and the immediate.
|
||||
pub const DRAWS_PER_INSTR: u64 = 4;
|
||||
/// The floor of chip operations per item (the x8 mixer with its constants hoisted: 72 x 128).
|
||||
pub const OPS_FLOOR_X8: u64 = 9_216;
|
||||
/// The floor of GPU operations per item (the x8 mixer as written: 72 x 144).
|
||||
pub const GPU_OPS_FLOOR_X8: u64 = 10_368;
|
||||
/// The floor of multiplies per item (the x8 mixer's 72 x 16).
|
||||
pub const MULS_FLOOR_X8: u64 = 1_152;
|
||||
/// Distinct rotation amounts an item's programs must use, at least.
|
||||
pub const DISTINCT_ROTS_FLOOR: usize = 8;
|
||||
/// Attempts before the generator gives up (never reached: see [`DeriveProgram::draw`]).
|
||||
pub const MAX_ATTEMPTS: u32 = 64;
|
||||
|
||||
/// The twelve forms. `c` is the chain register (the previous destination), `d` the destination (`d != c`), `b` a
|
||||
/// third register (`b != d`, `b != c`), `k` a rotation in 1..31, `i` a 32-bit constant (odd for `MulC`).
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
|
||||
#[repr(u8)]
|
||||
pub enum DOp {
|
||||
/// `d += c`
|
||||
Add = 0,
|
||||
/// `d -= c`
|
||||
Sub = 1,
|
||||
/// `d ^= c`
|
||||
Xor = 2,
|
||||
/// `d *= (c OR 1)`: the multiply-lo form, odd so it is a bijection on `d`
|
||||
Mul = 3,
|
||||
/// `d = rotl(d, k) + c`
|
||||
Rot = 4,
|
||||
/// `d = rotl(d ^ c, k)`
|
||||
XRot = 5,
|
||||
/// `d += c + i`
|
||||
AddC = 6,
|
||||
/// `d ^= c ^ i`
|
||||
XorC = 7,
|
||||
/// `d = (d ^ c) * i`, `i` odd: the per-word form of `M_r` with the chain in place of the round constant
|
||||
MulC = 8,
|
||||
/// `d = d * i + c`, `i` odd
|
||||
MulC2 = 9,
|
||||
/// `d ^= (c AND b)`
|
||||
AndX = 10,
|
||||
/// `d += (c OR b)`
|
||||
OrX = 11,
|
||||
}
|
||||
|
||||
/// The op weights in percent, in draw order (sum 100). Fixed at genesis; only the order, the registers and the
|
||||
/// constants are drawn.
|
||||
pub const DOP_WEIGHTS: [(DOp, u64); 12] = [
|
||||
(DOp::Add, 14),
|
||||
(DOp::Sub, 10),
|
||||
(DOp::Xor, 14),
|
||||
(DOp::Mul, 10),
|
||||
(DOp::Rot, 10),
|
||||
(DOp::XRot, 10),
|
||||
(DOp::AddC, 6),
|
||||
(DOp::XorC, 6),
|
||||
(DOp::MulC, 8),
|
||||
(DOp::MulC2, 4),
|
||||
(DOp::AndX, 4),
|
||||
(DOp::OrX, 4),
|
||||
];
|
||||
|
||||
impl DOp {
|
||||
pub fn from_u8(v: u8) -> Option<DOp> {
|
||||
DOP_WEIGHTS.iter().map(|(o, _)| *o).find(|o| *o as u8 == v)
|
||||
}
|
||||
pub fn name(self) -> &'static str {
|
||||
match self {
|
||||
DOp::Add => "add",
|
||||
DOp::Sub => "sub",
|
||||
DOp::Xor => "xor",
|
||||
DOp::Mul => "mul",
|
||||
DOp::Rot => "rot",
|
||||
DOp::XRot => "xrot",
|
||||
DOp::AddC => "addc",
|
||||
DOp::XorC => "xorc",
|
||||
DOp::MulC => "mulc",
|
||||
DOp::MulC2 => "mulc2",
|
||||
DOp::AndX => "andx",
|
||||
DOp::OrX => "orx",
|
||||
}
|
||||
}
|
||||
/// Integer operations as a GPU executes the form (every `|`, `&`, `+`, `^`, `*`, rotate counts one).
|
||||
pub fn gpu_ops(self) -> u64 {
|
||||
match self {
|
||||
DOp::Add | DOp::Sub | DOp::Xor => 1,
|
||||
_ => 2,
|
||||
}
|
||||
}
|
||||
/// Integer operations as the chip model counts them (`c OR 1` is a wire on a chip, so `Mul` is one multiply;
|
||||
/// a constant folded into a chain value is still an add or an xor, so every other two-op form stays two).
|
||||
pub fn chip_ops(self) -> u64 {
|
||||
match self {
|
||||
DOp::Add | DOp::Sub | DOp::Xor | DOp::Mul => 1,
|
||||
_ => 2,
|
||||
}
|
||||
}
|
||||
pub fn is_mul(self) -> bool {
|
||||
matches!(self, DOp::Mul | DOp::MulC | DOp::MulC2)
|
||||
}
|
||||
pub fn has_rot(self) -> bool {
|
||||
matches!(self, DOp::Rot | DOp::XRot)
|
||||
}
|
||||
pub fn has_third(self) -> bool {
|
||||
matches!(self, DOp::AndX | DOp::OrX)
|
||||
}
|
||||
pub fn has_imm(self) -> bool {
|
||||
matches!(self, DOp::AddC | DOp::XorC | DOp::MulC | DOp::MulC2)
|
||||
}
|
||||
/// The op of a roll in 0..99.
|
||||
pub fn for_roll(roll: u64) -> DOp {
|
||||
let mut acc = 0u64;
|
||||
for (op, w) in DOP_WEIGHTS {
|
||||
acc += w;
|
||||
if roll < acc {
|
||||
return op;
|
||||
}
|
||||
}
|
||||
DOp::OrX
|
||||
}
|
||||
}
|
||||
|
||||
/// One instruction. `src` is the chain register (carried so the interpreter and the emitter need no state).
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
|
||||
pub struct DInstr {
|
||||
pub op: DOp,
|
||||
pub dst: u8,
|
||||
pub src: u8,
|
||||
/// The third register of `AndX` and `OrX`; 0 on every other form (drawn and unused).
|
||||
pub src2: u8,
|
||||
/// The rotation 1..31 of `Rot` and `XRot`, else 0.
|
||||
pub rot: u8,
|
||||
/// The constant of `AddC`, `XorC` (any), `MulC` and `MulC2` (odd); 0 on every other form.
|
||||
pub imm: u32,
|
||||
}
|
||||
|
||||
/// The nine round programs of an item for one day, with the attempt that passed the acceptance test.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct DeriveProgram {
|
||||
pub len: u32,
|
||||
pub attempt: u32,
|
||||
pub rounds: Vec<Vec<DInstr>>,
|
||||
}
|
||||
|
||||
/// Why a candidate was rejected.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub enum DeriveReject {
|
||||
/// A register no instruction of round program `round` writes.
|
||||
RegisterNeverWritten { round: usize, reg: u8 },
|
||||
/// Fewer than [`DISTINCT_ROTS_FLOOR`] distinct rotation amounts over the item's programs.
|
||||
RotationsDegenerate { distinct: usize },
|
||||
/// Chip operations per item under [`OPS_FLOOR_X8`] scaled to the length.
|
||||
OpsUnderFloor { ops: u64, floor: u64 },
|
||||
/// GPU operations per item under [`GPU_OPS_FLOOR_X8`] scaled to the length.
|
||||
GpuOpsUnderFloor { ops: u64, floor: u64 },
|
||||
/// Multiplies per item under [`MULS_FLOOR_X8`] scaled to the length.
|
||||
MulsUnderFloor { muls: u64, floor: u64 },
|
||||
}
|
||||
|
||||
impl std::fmt::Display for DeriveReject {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
match self {
|
||||
DeriveReject::RegisterNeverWritten { round, reg } => write!(f, "register {reg} never written in round program {round}"),
|
||||
DeriveReject::RotationsDegenerate { distinct } => write!(f, "only {distinct} distinct rotation amounts"),
|
||||
DeriveReject::OpsUnderFloor { ops, floor } => write!(f, "{ops} chip operations per item, floor {floor}"),
|
||||
DeriveReject::GpuOpsUnderFloor { ops, floor } => write!(f, "{ops} GPU operations per item, floor {floor}"),
|
||||
DeriveReject::MulsUnderFloor { muls, floor } => write!(f, "{muls} multiplies per item, floor {floor}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl DeriveProgram {
|
||||
/// Draw one candidate of `len` instructions per round program from `rng` (four draws per instruction).
|
||||
pub fn draw_candidate(rng: &mut SplitMix64, len: u32, attempt: u32) -> DeriveProgram {
|
||||
let mut rounds = Vec::with_capacity(DERIVE_PROGRAMS);
|
||||
for _ in 0..DERIVE_PROGRAMS {
|
||||
let mut prog = Vec::with_capacity(len as usize);
|
||||
let mut chain = 0u8;
|
||||
for _ in 0..len {
|
||||
let op = DOp::for_roll(rng.below(100));
|
||||
// the destination: the 15 registers other than the chain, in ascending order
|
||||
let d_roll = rng.below((DERIVE_REGS - 1) as u64) as u8;
|
||||
let dst = if d_roll >= chain { d_roll + 1 } else { d_roll };
|
||||
// the third register: the 14 registers other than dst and the chain, in ascending order
|
||||
let b_roll = rng.below((DERIVE_REGS - 2) as u64) as u8;
|
||||
let (lo, hi) = if dst < chain { (dst, chain) } else { (chain, dst) };
|
||||
let mut b = b_roll;
|
||||
if b >= lo {
|
||||
b += 1;
|
||||
}
|
||||
if b >= hi {
|
||||
b += 1;
|
||||
}
|
||||
let x = rng.next() as u32;
|
||||
let mut ins = DInstr { op, dst, src: chain, src2: 0, rot: 0, imm: 0 };
|
||||
if op.has_third() {
|
||||
ins.src2 = b;
|
||||
}
|
||||
if op.has_rot() {
|
||||
ins.rot = 1 + (x % 31) as u8;
|
||||
}
|
||||
if op.has_imm() {
|
||||
ins.imm = if matches!(op, DOp::MulC | DOp::MulC2) { x | 1 } else { x };
|
||||
}
|
||||
prog.push(ins);
|
||||
chain = dst;
|
||||
}
|
||||
rounds.push(prog);
|
||||
}
|
||||
DeriveProgram { len, attempt, rounds }
|
||||
}
|
||||
|
||||
/// Draw the program of a day: candidates from `rng` in turn until one passes [`DeriveProgram::check`].
|
||||
/// Panics after [`MAX_ATTEMPTS`] (the floors sit more than 7 standard deviations under the expected counts, so
|
||||
/// a rejection is a rare event and 64 in a row is not one that happens).
|
||||
pub fn draw(rng: &mut SplitMix64, len: u32) -> DeriveProgram {
|
||||
for attempt in 0..MAX_ATTEMPTS {
|
||||
let p = Self::draw_candidate(rng, len, attempt);
|
||||
if p.check().is_ok() {
|
||||
return p;
|
||||
}
|
||||
}
|
||||
panic!("derivation program: {MAX_ATTEMPTS} candidates rejected in a row");
|
||||
}
|
||||
|
||||
/// The floors for this length (chip operations, GPU operations, multiplies): the x8 floors scaled by
|
||||
/// `len / DERIVE_LEN_X8`, so a shorter class, measured as a fallback, has its own proportional floors.
|
||||
pub fn floors(len: u32) -> (u64, u64, u64) {
|
||||
let scale = |f: u64| f * len as u64 / DERIVE_LEN_X8 as u64;
|
||||
(scale(OPS_FLOOR_X8), scale(GPU_OPS_FLOOR_X8), scale(MULS_FLOOR_X8))
|
||||
}
|
||||
|
||||
/// The acceptance test: every register written in every round program; at least [`DISTINCT_ROTS_FLOOR`]
|
||||
/// distinct rotation amounts; chip operations, GPU operations and multiplies per item at or above the floors
|
||||
/// (the x8 mixer's counts from the code). The structural
|
||||
/// rules (`dst != src`, the third register distinct, rotations in 1..31, odd multiplier constants) hold by
|
||||
/// construction and are asserted.
|
||||
pub fn check(&self) -> Result<(), DeriveReject> {
|
||||
let mut rots = [false; 32];
|
||||
for (r, prog) in self.rounds.iter().enumerate() {
|
||||
let mut written = [false; DERIVE_REGS];
|
||||
let mut chain = 0u8;
|
||||
for ins in prog {
|
||||
assert!(ins.src == chain && ins.dst != ins.src && (ins.dst as usize) < DERIVE_REGS, "chain rule");
|
||||
if ins.op.has_third() {
|
||||
assert!(ins.src2 != ins.dst && ins.src2 != ins.src && (ins.src2 as usize) < DERIVE_REGS, "third register");
|
||||
}
|
||||
if ins.op.has_rot() {
|
||||
assert!((1..=31).contains(&ins.rot), "rotation");
|
||||
rots[ins.rot as usize] = true;
|
||||
}
|
||||
if matches!(ins.op, DOp::MulC | DOp::MulC2) {
|
||||
assert!(ins.imm & 1 == 1, "odd multiplier");
|
||||
}
|
||||
written[ins.dst as usize] = true;
|
||||
chain = ins.dst;
|
||||
}
|
||||
if let Some(reg) = written.iter().position(|w| !w) {
|
||||
return Err(DeriveReject::RegisterNeverWritten { round: r, reg: reg as u8 });
|
||||
}
|
||||
}
|
||||
let distinct = rots.iter().filter(|r| **r).count();
|
||||
if distinct < DISTINCT_ROTS_FLOOR {
|
||||
return Err(DeriveReject::RotationsDegenerate { distinct });
|
||||
}
|
||||
let (ops_floor, gpu_floor, muls_floor) = Self::floors(self.len);
|
||||
let ops = self.chip_ops();
|
||||
if ops < ops_floor {
|
||||
return Err(DeriveReject::OpsUnderFloor { ops, floor: ops_floor });
|
||||
}
|
||||
let gpu = self.gpu_ops();
|
||||
if gpu < gpu_floor {
|
||||
return Err(DeriveReject::GpuOpsUnderFloor { ops: gpu, floor: gpu_floor });
|
||||
}
|
||||
let muls = self.muls();
|
||||
if muls < muls_floor {
|
||||
return Err(DeriveReject::MulsUnderFloor { muls, floor: muls_floor });
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
pub fn instr_count(&self) -> u64 {
|
||||
self.rounds.iter().map(|p| p.len() as u64).sum()
|
||||
}
|
||||
pub fn gpu_ops(&self) -> u64 {
|
||||
self.rounds.iter().flatten().map(|i| i.op.gpu_ops()).sum()
|
||||
}
|
||||
pub fn chip_ops(&self) -> u64 {
|
||||
self.rounds.iter().flatten().map(|i| i.op.chip_ops()).sum()
|
||||
}
|
||||
pub fn muls(&self) -> u64 {
|
||||
self.rounds.iter().flatten().filter(|i| i.op.is_mul()).count() as u64
|
||||
}
|
||||
/// Count per op, in [`DOP_WEIGHTS`] order.
|
||||
pub fn op_counts(&self) -> [u64; 12] {
|
||||
let mut c = [0u64; 12];
|
||||
for i in self.rounds.iter().flatten() {
|
||||
c[i.op as usize] += 1;
|
||||
}
|
||||
c
|
||||
}
|
||||
/// "add=887 sub=..." in weight order.
|
||||
pub fn op_mix(&self) -> String {
|
||||
let c = self.op_counts();
|
||||
DOP_WEIGHTS.iter().map(|(o, _)| format!("{}={}", o.name(), c[*o as usize])).collect::<Vec<_>>().join(" ")
|
||||
}
|
||||
/// FNV-1a 64 over the instruction stream (op, dst, src, src2, rot, imm as bytes): the program's fingerprint
|
||||
/// for packs and logs.
|
||||
pub fn fingerprint(&self) -> u64 {
|
||||
let mut b = Vec::with_capacity(self.instr_count() as usize * 9);
|
||||
for i in self.rounds.iter().flatten() {
|
||||
b.push(i.op as u8);
|
||||
b.push(i.dst);
|
||||
b.push(i.src);
|
||||
b.push(i.src2);
|
||||
b.push(i.rot);
|
||||
b.extend_from_slice(&i.imm.to_le_bytes());
|
||||
}
|
||||
crate::seed::fnv1a64(&b)
|
||||
}
|
||||
}
|
||||
|
||||
/// Lanes of the SoA interpreter: the verifier derives up to 32 distinct items per load (one per lane of the
|
||||
/// unit), so each instruction runs across 32 item states at once and the dispatch is paid once per 32 items.
|
||||
pub const SOA_LANES: usize = 32;
|
||||
|
||||
/// The item states of a batch, word-major: `st[reg][lane]`.
|
||||
pub type SoaState = [[u32; SOA_LANES]; DERIVE_REGS];
|
||||
|
||||
#[inline(always)]
|
||||
fn rotl(x: u32, n: u32) -> u32 {
|
||||
x.rotate_left(n)
|
||||
}
|
||||
|
||||
/// The twelve forms over a batch, one function each, every one a straight loop over the lanes the compiler
|
||||
/// vectorises. The destination row and the source rows are distinct by the chain rule (`dst != src`, and the third
|
||||
/// register distinct from both: asserted by [`DeriveProgram::check`] and checked here in debug builds), so the
|
||||
/// rows are addressed through raw pointers rather than copied out of the state.
|
||||
mod forms {
|
||||
use super::{rotl, DInstr, SoaState, SOA_LANES};
|
||||
#[inline(always)]
|
||||
pub fn add(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = d.wrapping_add(*c);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn sub(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = d.wrapping_sub(*c);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn xor(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d ^= *c;
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn mul(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = d.wrapping_mul(*c | 1);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn rot(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
let r = ins.rot as u32;
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = rotl(*d, r).wrapping_add(*c);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn xrot(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
let r = ins.rot as u32;
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = rotl(*d ^ *c, r);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn addc(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
let i = ins.imm;
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = d.wrapping_add(c.wrapping_add(i));
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn xorc(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
let i = ins.imm;
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d ^= *c ^ i;
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn mulc(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
let i = ins.imm;
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = (*d ^ *c).wrapping_mul(i);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn mulc2(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src);
|
||||
let i = ins.imm;
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
|
||||
*d = d.wrapping_mul(i).wrapping_add(*c);
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn andx(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src && ins.src2 != ins.dst && ins.src2 != ins.src);
|
||||
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
let bp = st.as_ptr().add(ins.src2 as usize) as *const u32;
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
let b = &*bp.add(k);
|
||||
*d ^= *c & *b;
|
||||
}
|
||||
}
|
||||
}
|
||||
#[inline(always)]
|
||||
pub fn orx(ins: &DInstr, st: &mut SoaState) {
|
||||
debug_assert!(ins.dst != ins.src && ins.src2 != ins.dst && ins.src2 != ins.src);
|
||||
|
||||
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
|
||||
// pointers stay inside `st`.
|
||||
unsafe {
|
||||
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
|
||||
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
|
||||
let bp = st.as_ptr().add(ins.src2 as usize) as *const u32;
|
||||
for k in 0..SOA_LANES {
|
||||
let d = &mut *dp.add(k);
|
||||
let c = &*cp.add(k);
|
||||
let b = &*bp.add(k);
|
||||
*d = d.wrapping_add(*c | *b);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// One instruction over the batch (the single dispatch; [`run_round`] dispatches on pairs).
|
||||
#[inline(always)]
|
||||
pub fn run_instr(ins: &DInstr, st: &mut SoaState) {
|
||||
match ins.op {
|
||||
DOp::Add => forms::add(ins, st),
|
||||
DOp::Sub => forms::sub(ins, st),
|
||||
DOp::Xor => forms::xor(ins, st),
|
||||
DOp::Mul => forms::mul(ins, st),
|
||||
DOp::Rot => forms::rot(ins, st),
|
||||
DOp::XRot => forms::xrot(ins, st),
|
||||
DOp::AddC => forms::addc(ins, st),
|
||||
DOp::XorC => forms::xorc(ins, st),
|
||||
DOp::MulC => forms::mulc(ins, st),
|
||||
DOp::MulC2 => forms::mulc2(ins, st),
|
||||
DOp::AndX => forms::andx(ins, st),
|
||||
DOp::OrX => forms::orx(ins, st),
|
||||
}
|
||||
}
|
||||
|
||||
/// Run one round program over the batch. The dispatch is on PAIRS of instructions (144 arms, one indirect branch
|
||||
/// per two instructions): the op sequence of a drawn program is random, so the branch predictor misses most
|
||||
/// dispatches, and the miss (about 3.7 ns of the 7.2 ns an instruction cost per batch on one M5 Max core, measured
|
||||
/// with `examples/derive_perf.rs` on 6 October 2026, a functional run) is paid once per pair instead of once per
|
||||
/// instruction. The result is bit for bit that of [`run_instr`] in sequence.
|
||||
#[inline(never)]
|
||||
pub fn run_round(prog: &[DInstr], st: &mut SoaState) {
|
||||
let mut it = prog.chunks_exact(2);
|
||||
for pair in &mut it {
|
||||
let (a, b) = (&pair[0], &pair[1]);
|
||||
match (a.op as u8) * 12 + b.op as u8 {
|
||||
0 => { forms::add(a, st); forms::add(b, st); }
|
||||
1 => { forms::add(a, st); forms::sub(b, st); }
|
||||
2 => { forms::add(a, st); forms::xor(b, st); }
|
||||
3 => { forms::add(a, st); forms::mul(b, st); }
|
||||
4 => { forms::add(a, st); forms::rot(b, st); }
|
||||
5 => { forms::add(a, st); forms::xrot(b, st); }
|
||||
6 => { forms::add(a, st); forms::addc(b, st); }
|
||||
7 => { forms::add(a, st); forms::xorc(b, st); }
|
||||
8 => { forms::add(a, st); forms::mulc(b, st); }
|
||||
9 => { forms::add(a, st); forms::mulc2(b, st); }
|
||||
10 => { forms::add(a, st); forms::andx(b, st); }
|
||||
11 => { forms::add(a, st); forms::orx(b, st); }
|
||||
12 => { forms::sub(a, st); forms::add(b, st); }
|
||||
13 => { forms::sub(a, st); forms::sub(b, st); }
|
||||
14 => { forms::sub(a, st); forms::xor(b, st); }
|
||||
15 => { forms::sub(a, st); forms::mul(b, st); }
|
||||
16 => { forms::sub(a, st); forms::rot(b, st); }
|
||||
17 => { forms::sub(a, st); forms::xrot(b, st); }
|
||||
18 => { forms::sub(a, st); forms::addc(b, st); }
|
||||
19 => { forms::sub(a, st); forms::xorc(b, st); }
|
||||
20 => { forms::sub(a, st); forms::mulc(b, st); }
|
||||
21 => { forms::sub(a, st); forms::mulc2(b, st); }
|
||||
22 => { forms::sub(a, st); forms::andx(b, st); }
|
||||
23 => { forms::sub(a, st); forms::orx(b, st); }
|
||||
24 => { forms::xor(a, st); forms::add(b, st); }
|
||||
25 => { forms::xor(a, st); forms::sub(b, st); }
|
||||
26 => { forms::xor(a, st); forms::xor(b, st); }
|
||||
27 => { forms::xor(a, st); forms::mul(b, st); }
|
||||
28 => { forms::xor(a, st); forms::rot(b, st); }
|
||||
29 => { forms::xor(a, st); forms::xrot(b, st); }
|
||||
30 => { forms::xor(a, st); forms::addc(b, st); }
|
||||
31 => { forms::xor(a, st); forms::xorc(b, st); }
|
||||
32 => { forms::xor(a, st); forms::mulc(b, st); }
|
||||
33 => { forms::xor(a, st); forms::mulc2(b, st); }
|
||||
34 => { forms::xor(a, st); forms::andx(b, st); }
|
||||
35 => { forms::xor(a, st); forms::orx(b, st); }
|
||||
36 => { forms::mul(a, st); forms::add(b, st); }
|
||||
37 => { forms::mul(a, st); forms::sub(b, st); }
|
||||
38 => { forms::mul(a, st); forms::xor(b, st); }
|
||||
39 => { forms::mul(a, st); forms::mul(b, st); }
|
||||
40 => { forms::mul(a, st); forms::rot(b, st); }
|
||||
41 => { forms::mul(a, st); forms::xrot(b, st); }
|
||||
42 => { forms::mul(a, st); forms::addc(b, st); }
|
||||
43 => { forms::mul(a, st); forms::xorc(b, st); }
|
||||
44 => { forms::mul(a, st); forms::mulc(b, st); }
|
||||
45 => { forms::mul(a, st); forms::mulc2(b, st); }
|
||||
46 => { forms::mul(a, st); forms::andx(b, st); }
|
||||
47 => { forms::mul(a, st); forms::orx(b, st); }
|
||||
48 => { forms::rot(a, st); forms::add(b, st); }
|
||||
49 => { forms::rot(a, st); forms::sub(b, st); }
|
||||
50 => { forms::rot(a, st); forms::xor(b, st); }
|
||||
51 => { forms::rot(a, st); forms::mul(b, st); }
|
||||
52 => { forms::rot(a, st); forms::rot(b, st); }
|
||||
53 => { forms::rot(a, st); forms::xrot(b, st); }
|
||||
54 => { forms::rot(a, st); forms::addc(b, st); }
|
||||
55 => { forms::rot(a, st); forms::xorc(b, st); }
|
||||
56 => { forms::rot(a, st); forms::mulc(b, st); }
|
||||
57 => { forms::rot(a, st); forms::mulc2(b, st); }
|
||||
58 => { forms::rot(a, st); forms::andx(b, st); }
|
||||
59 => { forms::rot(a, st); forms::orx(b, st); }
|
||||
60 => { forms::xrot(a, st); forms::add(b, st); }
|
||||
61 => { forms::xrot(a, st); forms::sub(b, st); }
|
||||
62 => { forms::xrot(a, st); forms::xor(b, st); }
|
||||
63 => { forms::xrot(a, st); forms::mul(b, st); }
|
||||
64 => { forms::xrot(a, st); forms::rot(b, st); }
|
||||
65 => { forms::xrot(a, st); forms::xrot(b, st); }
|
||||
66 => { forms::xrot(a, st); forms::addc(b, st); }
|
||||
67 => { forms::xrot(a, st); forms::xorc(b, st); }
|
||||
68 => { forms::xrot(a, st); forms::mulc(b, st); }
|
||||
69 => { forms::xrot(a, st); forms::mulc2(b, st); }
|
||||
70 => { forms::xrot(a, st); forms::andx(b, st); }
|
||||
71 => { forms::xrot(a, st); forms::orx(b, st); }
|
||||
72 => { forms::addc(a, st); forms::add(b, st); }
|
||||
73 => { forms::addc(a, st); forms::sub(b, st); }
|
||||
74 => { forms::addc(a, st); forms::xor(b, st); }
|
||||
75 => { forms::addc(a, st); forms::mul(b, st); }
|
||||
76 => { forms::addc(a, st); forms::rot(b, st); }
|
||||
77 => { forms::addc(a, st); forms::xrot(b, st); }
|
||||
78 => { forms::addc(a, st); forms::addc(b, st); }
|
||||
79 => { forms::addc(a, st); forms::xorc(b, st); }
|
||||
80 => { forms::addc(a, st); forms::mulc(b, st); }
|
||||
81 => { forms::addc(a, st); forms::mulc2(b, st); }
|
||||
82 => { forms::addc(a, st); forms::andx(b, st); }
|
||||
83 => { forms::addc(a, st); forms::orx(b, st); }
|
||||
84 => { forms::xorc(a, st); forms::add(b, st); }
|
||||
85 => { forms::xorc(a, st); forms::sub(b, st); }
|
||||
86 => { forms::xorc(a, st); forms::xor(b, st); }
|
||||
87 => { forms::xorc(a, st); forms::mul(b, st); }
|
||||
88 => { forms::xorc(a, st); forms::rot(b, st); }
|
||||
89 => { forms::xorc(a, st); forms::xrot(b, st); }
|
||||
90 => { forms::xorc(a, st); forms::addc(b, st); }
|
||||
91 => { forms::xorc(a, st); forms::xorc(b, st); }
|
||||
92 => { forms::xorc(a, st); forms::mulc(b, st); }
|
||||
93 => { forms::xorc(a, st); forms::mulc2(b, st); }
|
||||
94 => { forms::xorc(a, st); forms::andx(b, st); }
|
||||
95 => { forms::xorc(a, st); forms::orx(b, st); }
|
||||
96 => { forms::mulc(a, st); forms::add(b, st); }
|
||||
97 => { forms::mulc(a, st); forms::sub(b, st); }
|
||||
98 => { forms::mulc(a, st); forms::xor(b, st); }
|
||||
99 => { forms::mulc(a, st); forms::mul(b, st); }
|
||||
100 => { forms::mulc(a, st); forms::rot(b, st); }
|
||||
101 => { forms::mulc(a, st); forms::xrot(b, st); }
|
||||
102 => { forms::mulc(a, st); forms::addc(b, st); }
|
||||
103 => { forms::mulc(a, st); forms::xorc(b, st); }
|
||||
104 => { forms::mulc(a, st); forms::mulc(b, st); }
|
||||
105 => { forms::mulc(a, st); forms::mulc2(b, st); }
|
||||
106 => { forms::mulc(a, st); forms::andx(b, st); }
|
||||
107 => { forms::mulc(a, st); forms::orx(b, st); }
|
||||
108 => { forms::mulc2(a, st); forms::add(b, st); }
|
||||
109 => { forms::mulc2(a, st); forms::sub(b, st); }
|
||||
110 => { forms::mulc2(a, st); forms::xor(b, st); }
|
||||
111 => { forms::mulc2(a, st); forms::mul(b, st); }
|
||||
112 => { forms::mulc2(a, st); forms::rot(b, st); }
|
||||
113 => { forms::mulc2(a, st); forms::xrot(b, st); }
|
||||
114 => { forms::mulc2(a, st); forms::addc(b, st); }
|
||||
115 => { forms::mulc2(a, st); forms::xorc(b, st); }
|
||||
116 => { forms::mulc2(a, st); forms::mulc(b, st); }
|
||||
117 => { forms::mulc2(a, st); forms::mulc2(b, st); }
|
||||
118 => { forms::mulc2(a, st); forms::andx(b, st); }
|
||||
119 => { forms::mulc2(a, st); forms::orx(b, st); }
|
||||
120 => { forms::andx(a, st); forms::add(b, st); }
|
||||
121 => { forms::andx(a, st); forms::sub(b, st); }
|
||||
122 => { forms::andx(a, st); forms::xor(b, st); }
|
||||
123 => { forms::andx(a, st); forms::mul(b, st); }
|
||||
124 => { forms::andx(a, st); forms::rot(b, st); }
|
||||
125 => { forms::andx(a, st); forms::xrot(b, st); }
|
||||
126 => { forms::andx(a, st); forms::addc(b, st); }
|
||||
127 => { forms::andx(a, st); forms::xorc(b, st); }
|
||||
128 => { forms::andx(a, st); forms::mulc(b, st); }
|
||||
129 => { forms::andx(a, st); forms::mulc2(b, st); }
|
||||
130 => { forms::andx(a, st); forms::andx(b, st); }
|
||||
131 => { forms::andx(a, st); forms::orx(b, st); }
|
||||
132 => { forms::orx(a, st); forms::add(b, st); }
|
||||
133 => { forms::orx(a, st); forms::sub(b, st); }
|
||||
134 => { forms::orx(a, st); forms::xor(b, st); }
|
||||
135 => { forms::orx(a, st); forms::mul(b, st); }
|
||||
136 => { forms::orx(a, st); forms::rot(b, st); }
|
||||
137 => { forms::orx(a, st); forms::xrot(b, st); }
|
||||
138 => { forms::orx(a, st); forms::addc(b, st); }
|
||||
139 => { forms::orx(a, st); forms::xorc(b, st); }
|
||||
140 => { forms::orx(a, st); forms::mulc(b, st); }
|
||||
141 => { forms::orx(a, st); forms::mulc2(b, st); }
|
||||
142 => { forms::orx(a, st); forms::andx(b, st); }
|
||||
143 => { forms::orx(a, st); forms::orx(b, st); }
|
||||
_ => unreachable!(),
|
||||
}
|
||||
}
|
||||
for ins in it.remainder() {
|
||||
run_instr(ins, st);
|
||||
}
|
||||
}
|
||||
|
||||
/// The scalar reference: one instruction on one 16-word state, the text the kernels carry (`emit.rs`,
|
||||
/// `derive_instr_text`) restated in Rust. The tests pin the SoA interpreter against it.
|
||||
pub fn run_round_scalar(prog: &[DInstr], s: &mut [u32; DERIVE_REGS]) {
|
||||
for ins in prog {
|
||||
let d = ins.dst as usize;
|
||||
let c = s[ins.src as usize];
|
||||
match ins.op {
|
||||
DOp::Add => s[d] = s[d].wrapping_add(c),
|
||||
DOp::Sub => s[d] = s[d].wrapping_sub(c),
|
||||
DOp::Xor => s[d] ^= c,
|
||||
DOp::Mul => s[d] = s[d].wrapping_mul(c | 1),
|
||||
DOp::Rot => s[d] = rotl(s[d], ins.rot as u32).wrapping_add(c),
|
||||
DOp::XRot => s[d] = rotl(s[d] ^ c, ins.rot as u32),
|
||||
DOp::AddC => s[d] = s[d].wrapping_add(c.wrapping_add(ins.imm)),
|
||||
DOp::XorC => s[d] ^= c ^ ins.imm,
|
||||
DOp::MulC => s[d] = (s[d] ^ c).wrapping_mul(ins.imm),
|
||||
DOp::MulC2 => s[d] = s[d].wrapping_mul(ins.imm).wrapping_add(c),
|
||||
DOp::AndX => s[d] ^= c & s[ins.src2 as usize],
|
||||
DOp::OrX => s[d] = s[d].wrapping_add(c | s[ins.src2 as usize]),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The source text of one instruction in the C-family dialects (the same text in Metal, CUDA C and OpenCL C:
|
||||
/// `s` is the 16-word state, `mh_rotl` the rotate of the memhard core).
|
||||
pub fn instr_text(ins: &DInstr) -> String {
|
||||
let (d, c, b) = (ins.dst, ins.src, ins.src2);
|
||||
match ins.op {
|
||||
DOp::Add => format!("s[{d}] += s[{c}];"),
|
||||
DOp::Sub => format!("s[{d}] -= s[{c}];"),
|
||||
DOp::Xor => format!("s[{d}] ^= s[{c}];"),
|
||||
DOp::Mul => format!("s[{d}] *= (s[{c}] | 1u);"),
|
||||
DOp::Rot => format!("s[{d}] = mh_rotl(s[{d}], {}u) + s[{c}];", ins.rot),
|
||||
DOp::XRot => format!("s[{d}] = mh_rotl(s[{d}] ^ s[{c}], {}u);", ins.rot),
|
||||
DOp::AddC => format!("s[{d}] += s[{c}] + {:#010x}u;", ins.imm),
|
||||
DOp::XorC => format!("s[{d}] ^= s[{c}] ^ {:#010x}u;", ins.imm),
|
||||
DOp::MulC => format!("s[{d}] = (s[{d}] ^ s[{c}]) * {:#010x}u;", ins.imm),
|
||||
DOp::MulC2 => format!("s[{d}] = s[{d}] * {:#010x}u + s[{c}];", ins.imm),
|
||||
DOp::AndX => format!("s[{d}] ^= (s[{c}] & s[{b}]);"),
|
||||
DOp::OrX => format!("s[{d}] += (s[{c}] | s[{b}]);"),
|
||||
}
|
||||
}
|
||||
|
||||
/// One instruction as a line of program.json: `"add d=3 c=0"`, `"mulc d=5 c=3 imm=0x..."`.
|
||||
pub fn instr_line(ins: &DInstr) -> String {
|
||||
let mut s = format!("{} d={} c={}", ins.op.name(), ins.dst, ins.src);
|
||||
if ins.op.has_third() {
|
||||
s.push_str(&format!(" b={}", ins.src2));
|
||||
}
|
||||
if ins.op.has_rot() {
|
||||
s.push_str(&format!(" k={}", ins.rot));
|
||||
}
|
||||
if ins.op.has_imm() {
|
||||
s.push_str(&format!(" imm={:#010x}", ins.imm));
|
||||
}
|
||||
s
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn weights_sum_and_ops() {
|
||||
assert_eq!(DOP_WEIGHTS.iter().map(|(_, w)| w).sum::<u64>(), 100);
|
||||
for (i, (op, _)) in DOP_WEIGHTS.iter().enumerate() {
|
||||
assert_eq!(*op as usize, i);
|
||||
assert_eq!(DOp::from_u8(i as u8), Some(*op));
|
||||
}
|
||||
assert_eq!(DOp::for_roll(0), DOp::Add);
|
||||
assert_eq!(DOp::for_roll(13), DOp::Add);
|
||||
assert_eq!(DOp::for_roll(14), DOp::Sub);
|
||||
assert_eq!(DOp::for_roll(99), DOp::OrX);
|
||||
// the expected chip operations per instruction, 1.52 (1.62 on a GPU), put 736 x 9 over the x8 floors
|
||||
let mean: f64 = DOP_WEIGHTS.iter().map(|(o, w)| o.chip_ops() as f64 * *w as f64 / 100.0).sum();
|
||||
assert!((mean - 1.52).abs() < 1e-9, "{mean}");
|
||||
let gpu_mean: f64 = DOP_WEIGHTS.iter().map(|(o, w)| o.gpu_ops() as f64 * *w as f64 / 100.0).sum();
|
||||
assert!((gpu_mean - 1.62).abs() < 1e-9, "{gpu_mean}");
|
||||
assert!(9.0 * DERIVE_LEN_X8 as f64 * mean > OPS_FLOOR_X8 as f64);
|
||||
assert!(9.0 * DERIVE_LEN_X8 as f64 * gpu_mean > GPU_OPS_FLOOR_X8 as f64);
|
||||
assert_eq!(DeriveProgram::floors(DERIVE_LEN_X8), (9_216, 10_368, 1_152));
|
||||
assert_eq!(DeriveProgram::floors(368), (4_608, 5_184, 576));
|
||||
let mul_share: f64 = DOP_WEIGHTS.iter().filter(|(o, _)| o.is_mul()).map(|(_, w)| *w as f64 / 100.0).sum();
|
||||
assert!(9.0 * DERIVE_LEN_X8 as f64 * mul_share > MULS_FLOOR_X8 as f64);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn draw_is_structural_and_accepted() {
|
||||
let mut rng = SplitMix64::new(0x1234_5678_9abc_def0);
|
||||
let p = DeriveProgram::draw(&mut rng, DERIVE_LEN_X8);
|
||||
assert_eq!(p.attempt, 0, "the first candidate of this seed passes");
|
||||
assert_eq!(p.rounds.len(), DERIVE_PROGRAMS);
|
||||
assert_eq!(p.instr_count(), 9 * DERIVE_LEN_X8 as u64);
|
||||
assert!(p.check().is_ok());
|
||||
assert!(p.chip_ops() >= OPS_FLOOR_X8 && p.gpu_ops() >= GPU_OPS_FLOOR_X8 && p.muls() >= MULS_FLOOR_X8);
|
||||
assert!(p.gpu_ops() > p.chip_ops());
|
||||
// every instruction consumes the newest result
|
||||
for prog in &p.rounds {
|
||||
let mut chain = 0u8;
|
||||
for ins in prog {
|
||||
assert_eq!(ins.src, chain);
|
||||
assert_ne!(ins.dst, chain);
|
||||
chain = ins.dst;
|
||||
}
|
||||
}
|
||||
// four draws per instruction: the same program again from the same seed, and a different one one draw on
|
||||
let mut rng2 = SplitMix64::new(0x1234_5678_9abc_def0);
|
||||
assert_eq!(DeriveProgram::draw(&mut rng2, DERIVE_LEN_X8), p);
|
||||
let mut rng3 = SplitMix64::new(0x1234_5678_9abc_def0);
|
||||
rng3.next();
|
||||
assert_ne!(DeriveProgram::draw(&mut rng3, DERIVE_LEN_X8), p);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn soa_matches_scalar_and_is_a_bijection() {
|
||||
let mut rng = SplitMix64::new(7);
|
||||
let p = DeriveProgram::draw(&mut rng, 64);
|
||||
let mut st: SoaState = [[0u32; SOA_LANES]; DERIVE_REGS];
|
||||
let mut scalars = [[0u32; DERIVE_REGS]; SOA_LANES];
|
||||
let mut x = SplitMix64::new(99);
|
||||
for k in 0..SOA_LANES {
|
||||
for r in 0..DERIVE_REGS {
|
||||
let v = x.next() as u32;
|
||||
st[r][k] = v;
|
||||
scalars[k][r] = v;
|
||||
}
|
||||
}
|
||||
let before = scalars;
|
||||
for prog in &p.rounds {
|
||||
run_round(prog, &mut st);
|
||||
for k in 0..SOA_LANES {
|
||||
run_round_scalar(prog, &mut scalars[k]);
|
||||
}
|
||||
}
|
||||
for k in 0..SOA_LANES {
|
||||
for r in 0..DERIVE_REGS {
|
||||
assert_eq!(st[r][k], scalars[k][r], "lane {k} reg {r}");
|
||||
}
|
||||
}
|
||||
// distinct inputs stay distinct (a bijection on the state, spot-checked: 32 lanes, no collision)
|
||||
for a in 0..SOA_LANES {
|
||||
for b in a + 1..SOA_LANES {
|
||||
assert_ne!(scalars[a], scalars[b]);
|
||||
assert_ne!(before[a], before[b]);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn acceptance_rejects_degenerate_draws() {
|
||||
let mut rng = SplitMix64::new(3);
|
||||
let mut p = DeriveProgram::draw(&mut rng, 64);
|
||||
// a register never written: make every write of round 2 go to the chain's neighbour
|
||||
let mut q = p.clone();
|
||||
for ins in q.rounds[2].iter_mut() {
|
||||
ins.dst = if ins.src == 1 { 2 } else { 1 };
|
||||
}
|
||||
let mut chain = 0u8;
|
||||
for ins in q.rounds[2].iter_mut() {
|
||||
ins.src = chain;
|
||||
ins.dst = if chain == 1 { 2 } else { 1 };
|
||||
chain = ins.dst;
|
||||
}
|
||||
assert!(matches!(q.check(), Err(DeriveReject::RegisterNeverWritten { round: 2, .. })));
|
||||
// all rotations equal
|
||||
let mut q = p.clone();
|
||||
for ins in q.rounds.iter_mut().flatten() {
|
||||
if ins.op.has_rot() {
|
||||
ins.rot = 5;
|
||||
}
|
||||
}
|
||||
assert!(matches!(q.check(), Err(DeriveReject::RotationsDegenerate { distinct: 1 })));
|
||||
// every op an add apart from the xor-rotates (so the rotations stay distinct): under the ops floor
|
||||
for ins in p.rounds.iter_mut().flatten() {
|
||||
if ins.op != DOp::XRot {
|
||||
ins.op = DOp::Add;
|
||||
ins.rot = 0;
|
||||
ins.imm = 0;
|
||||
ins.src2 = 0;
|
||||
}
|
||||
}
|
||||
assert!(matches!(p.check(), Err(DeriveReject::OpsUnderFloor { .. })), "{:?}", p.check());
|
||||
// no multiplies at all but the ops floor met: under the multiply floor
|
||||
let mut q = DeriveProgram::draw(&mut SplitMix64::new(11), 64);
|
||||
for ins in q.rounds.iter_mut().flatten() {
|
||||
if ins.op.is_mul() {
|
||||
ins.op = DOp::AddC;
|
||||
}
|
||||
}
|
||||
assert!(matches!(q.check(), Err(DeriveReject::MulsUnderFloor { .. })), "{:?}", q.check());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn text_forms() {
|
||||
let i = DInstr { op: DOp::MulC, dst: 5, src: 3, src2: 0, rot: 0, imm: 0x9e37_79b9 };
|
||||
assert_eq!(instr_text(&i), "s[5] = (s[5] ^ s[3]) * 0x9e3779b9u;");
|
||||
assert_eq!(instr_line(&i), "mulc d=5 c=3 imm=0x9e3779b9");
|
||||
let i = DInstr { op: DOp::XRot, dst: 0, src: 15, src2: 0, rot: 17, imm: 0 };
|
||||
assert_eq!(instr_text(&i), "s[0] = mh_rotl(s[0] ^ s[15], 17u);");
|
||||
let i = DInstr { op: DOp::AndX, dst: 2, src: 9, src2: 14, rot: 0, imm: 0 };
|
||||
assert_eq!(instr_text(&i), "s[2] ^= (s[9] & s[14]);");
|
||||
}
|
||||
}
|
||||
|
|
@ -10,6 +10,7 @@
|
|||
//! as a bare `0x003fffff`. The Swift writes it quoted (`jhex`), which is not valid JSON.
|
||||
|
||||
use crate::generator::{EraParams, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
|
||||
use crate::derive::{instr_line as derive_instr_line, instr_text as derive_instr_text, DERIVE_PROGRAMS};
|
||||
use crate::memhard::{
|
||||
hot_key, hot_segments, hot_words, Layout, MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG,
|
||||
CHACHA_ROUNDS, CHACHA_SIGMA, HOT_TAG, ITEM_ROUNDS,
|
||||
|
|
@ -195,6 +196,14 @@ fn class_header_lines(p: &Program) -> String {
|
|||
", p.class.mixer_mult));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_GROWTH {} // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
", p.class.growth as u8));
|
||||
}
|
||||
if p.class.derive_len != 0 {
|
||||
s.push_str("// Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md, a prototype, NOT class v3): the item
|
||||
");
|
||||
s.push_str("// derivation runs the day's drawn program (memhard.h: mh_round_0..8, IGNEUM_DERIVE_LEN instructions each) in place of the mixer.
|
||||
");
|
||||
s.push_str(&format!("#define IGNEUM_CLASS_DERIVE_LEN {}
|
||||
", p.class.derive_len));
|
||||
}
|
||||
s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {}
|
||||
", p.class.load_slots));
|
||||
|
|
@ -474,12 +483,18 @@ pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: La
|
|||
CACHE_LINES_PER_SEGMENT,
|
||||
CHACHA_ROUNDS
|
||||
));
|
||||
if m == 1 {
|
||||
if mp.derive.is_some() {
|
||||
s.push_str("// Item: 8 rounds of (the day's round program + one 64-byte cache read), then round program 8. All parameters are literals.\n");
|
||||
} else if m == 1 {
|
||||
s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n");
|
||||
} else {
|
||||
s.push_str(&format!("// Item: 8 rounds of {m} x seed-parameterised mixer + one 64-byte cache read, then {m} x final mixer (class v3, mixer multiplier {m},\n"));
|
||||
s.push_str("// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.\n");
|
||||
}
|
||||
if let Some(dp) = &mp.derive {
|
||||
s.push_str(&format!("// Counter ASIC 3.0 item 2 (docs/plans/counter-asic-3-derivation.md): the mixer slots run the day's drawn program, {} instructions per round\n", dp.len));
|
||||
s.push_str(&format!("// program (mh_round_0..{}), {} per item, drawn from the day key stream after the mixer constants (attempt {}, fingerprint {:016x}).\n", DERIVE_PROGRAMS - 1, dp.instr_count(), dp.attempt, dp.fingerprint()));
|
||||
}
|
||||
s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(cache_line_mask)));
|
||||
s.push_str(&format!("#define MH_SEGMENT_LINES {}u\n", CACHE_LINES_PER_SEGMENT));
|
||||
s.push_str("#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }\n");
|
||||
|
|
@ -559,6 +574,34 @@ pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: La
|
|||
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of {m} x mixer + cache line s[0] & mask; {m} x final mixer.\n"
|
||||
));
|
||||
}
|
||||
if let Some(dp) = &mp.derive {
|
||||
// the nine round programs as straight-line functions over the 16-word state (the same text in every dialect)
|
||||
for (r, prog) in dp.rounds.iter().enumerate() {
|
||||
s.push_str(&format!("// Round program {r}: {} instructions, chain rule (every instruction reads the register the previous one wrote; s[0] first).\n", prog.len()));
|
||||
s.push_str(&format!("{fn_} void mh_round_{r}({lptr} s) {{\n"));
|
||||
for ins in prog {
|
||||
s.push_str(" ");
|
||||
s.push_str(&derive_instr_text(ins));
|
||||
s.push('\n');
|
||||
}
|
||||
s.push_str("}\n");
|
||||
}
|
||||
s.push_str(&format!("// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of (round program r, cache line s[0] & mask); round program {ITEM_ROUNDS}.\n"));
|
||||
s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n"));
|
||||
for i in 0..8 {
|
||||
s.push_str(&format!(" s[{i}] = {};\n", hex(k[i])));
|
||||
}
|
||||
for i in 0..8 {
|
||||
s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(mul[i]), hex(c[i])));
|
||||
}
|
||||
for r in 0..ITEM_ROUNDS {
|
||||
s.push_str(&format!(" mh_round_{r}(s);\n"));
|
||||
s.push_str(&format!(" {{ {cptr} line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); for ({u} i = 0u; i < 16u; ++i) s[i] ^= line[i]; }}\n"));
|
||||
}
|
||||
s.push_str(&format!(" mh_round_{ITEM_ROUNDS}(s);\n"));
|
||||
s.push_str("}\n");
|
||||
return finish_memhard_core(s, layout, u, fn_, cptr);
|
||||
}
|
||||
s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n"));
|
||||
for i in 0..8 {
|
||||
s.push_str(&format!(" s[{i}] = {};\n", hex(k[i])));
|
||||
|
|
@ -584,6 +627,11 @@ pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: La
|
|||
));
|
||||
}
|
||||
s.push_str("}\n");
|
||||
finish_memhard_core(s, layout, u, fn_, cptr)
|
||||
}
|
||||
|
||||
/// The tail of the memhard core: `mh_word` (and the era layout helpers) after `mh_item`.
|
||||
fn finish_memhard_core(mut s: String, layout: Layout, u: &str, fn_: &str, cptr: &str) -> String {
|
||||
if layout.is_linear() {
|
||||
s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n");
|
||||
s.push_str(&format!(
|
||||
|
|
@ -1486,6 +1534,15 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
if mp.shape.mixer_mult != 1 {
|
||||
s.push_str(&format!("#define IGNEUM_MIXER_MULT {} // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)\n", mp.shape.mixer_mult));
|
||||
}
|
||||
if let Some(dp) = &mp.derive {
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_LEN {} // instructions per round program of the day's item-derivation program (Counter ASIC 3.0 item 2; memhard.h mh_round_0..8)\n", dp.len));
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_ATTEMPT {}\n", dp.attempt));
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_FINGERPRINT {}\n", hex64(dp.fingerprint())));
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_INSTRS_PER_ITEM {}\n", dp.instr_count()));
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_GPU_OPS_PER_ITEM {}\n", dp.gpu_ops()));
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_CHIP_OPS_PER_ITEM {}\n", dp.chip_ops()));
|
||||
s.push_str(&format!("#define IGNEUM_DERIVE_MULS_PER_ITEM {}\n", dp.muls()));
|
||||
}
|
||||
s.push_str(&format!(
|
||||
"#define IGNEUM_MIX_ROT_INIT {{ {} }}\n",
|
||||
mp.rot.iter().map(|r| format!("{r}u")).collect::<Vec<_>>().join(", ")
|
||||
|
|
@ -1694,6 +1751,10 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
if !p.class.is_v2() {
|
||||
let c = p.width_counts();
|
||||
s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name())));
|
||||
if p.class.derive_len != 0 {
|
||||
s.push_str(&format!(" \"derive_len\": {},\n", p.class.derive_len));
|
||||
s.push_str(" \"derive\": \"Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md; a prototype, not class v3): the nine mixer slots of the item derivation each run a straight-line program of derive_len instructions drawn from the day key stream after the 40 mixer draws, four draws per instruction (op roll below(100), destination roll below(15), third-register roll below(14), the immediate next()); every instruction reads the register the previous one wrote (s[0] first) and writes another; twelve forms, each a bijection on the state; the 8 dependent cache reads per item unchanged; the acceptance test of derive.rs (every register written per round program, 8 distinct rotations, the x8 mixer's operation and multiply counts as floors) rejects a draw and the next attempt continues the stream\",\n");
|
||||
}
|
||||
if p.class.mixer_mult != 1 || p.class.growth {
|
||||
s.push_str(&format!(" \"mixer_mult\": {},\n", p.class.mixer_mult));
|
||||
s.push_str(&format!(" \"cache_growth\": {},\n", p.class.growth));
|
||||
|
|
@ -1801,7 +1862,24 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
join_jhex(&mp.rc)
|
||||
));
|
||||
// The Swift writes jhex(cacheLineMask) here, which breaks the JSON. We write the bare literal.
|
||||
if shape.mixer_mult == 1 {
|
||||
if let Some(dp) = &mp.derive {
|
||||
s.push_str(&format!(
|
||||
" \"derive_len\": {},\n \"derive_attempt\": {},\n \"derive_fingerprint\": {},\n \"derive_op_mix\": {},\n \"derive_instrs_per_item\": {},\n \"derive_gpu_ops_per_item\": {},\n \"derive_chip_ops_per_item\": {},\n \"derive_muls_per_item\": {},\n",
|
||||
dp.len, dp.attempt, jhex64(dp.fingerprint()), jstr(&dp.op_mix()), dp.instr_count(), dp.gpu_ops(), dp.chip_ops(), dp.muls()
|
||||
));
|
||||
s.push_str(&format!(
|
||||
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = P_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = P_{ITEM_ROUNDS}(s); item(t) = s; P_r is round program r below (c = the register the previous instruction wrote, s[0] first): add d += c; sub d -= c; xor d ^= c; mul d *= (c | 1); rot d = rotl(d, k) + c; xrot d = rotl(d ^ c, k); addc d += c + imm; xorc d ^= c ^ imm; mulc d = (d ^ c) * imm; mulc2 d = d * imm + c; andx d ^= (c & b); orx d += (c | b)\",\n",
|
||||
ITEM_ROUNDS - 1,
|
||||
shape.cache_line_mask()
|
||||
));
|
||||
s.push_str(" \"programs\": [\n");
|
||||
for (r, prog) in dp.rounds.iter().enumerate() {
|
||||
s.push_str(" [");
|
||||
s.push_str(&prog.iter().map(|i| jstr(&derive_instr_line(i))).collect::<Vec<_>>().join(", "));
|
||||
s.push_str(if r + 1 < dp.rounds.len() { "],\n" } else { "]\n" });
|
||||
}
|
||||
s.push_str(" ],\n");
|
||||
} else if shape.mixer_mult == 1 {
|
||||
s.push_str(&format!(
|
||||
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n",
|
||||
ITEM_ROUNDS - 1,
|
||||
|
|
|
|||
|
|
@ -216,6 +216,10 @@ pub struct LoadClass {
|
|||
/// Hot table (`docs/plans/hot-table.md`, measured 5 October 2026 and not adopted): `Some(HotClass { mb, k, added })`
|
||||
/// turns `k` load slots into reads of an `mb` MiB epoch table. `None` for every other class, class v3 included.
|
||||
pub hot: Option<HotClass>,
|
||||
/// Counter ASIC 3.0 item 2 (`crate::derive`, `docs/plans/counter-asic-3-derivation.md`, 6 October 2026, a
|
||||
/// prototype behind the class): instructions per round program of the per-day item-derivation program that
|
||||
/// replaces the fixed mixer when non-zero (the mixer multiplier is then unused and 1). 0 for every other class.
|
||||
pub derive_len: u16,
|
||||
}
|
||||
|
||||
/// The parameters one era draws from its seed `E_n` (`docs/plans/era-layout.md` section 1.1, the proposed text of
|
||||
|
|
@ -365,13 +369,13 @@ impl LoadClass {
|
|||
impl LoadClass {
|
||||
/// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash.
|
||||
pub const V2: LoadClass =
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false, era: None, hot: None };
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false, era: None, hot: None, derive_len: 0 };
|
||||
|
||||
/// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`):
|
||||
/// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the
|
||||
/// mixer applied 4 times per round, and the cache growth rule. Name "mx4".
|
||||
pub const MX4: LoadClass =
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true, era: None, hot: None };
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true, era: None, hot: None, derive_len: 0 };
|
||||
|
||||
/// The era class over `base` (`docs/plans/era-layout.md`): the parameters drawn by [`era_draw`]; when `allowed`
|
||||
/// has more than one width the drawn width becomes the class mix (every load that width), otherwise the base
|
||||
|
|
@ -406,6 +410,24 @@ impl LoadClass {
|
|||
pub const MX8: LoadClass =
|
||||
LoadClass { mixer_mult: 8, ..LoadClass::MX4 };
|
||||
|
||||
/// Counter ASIC 3.0 item 2 (6 October 2026, `docs/plans/counter-asic-3-derivation.md`): version 2 loads and the
|
||||
/// growth rule of class v3, with the per-day derivation program of `crate::derive` (736 instructions per round
|
||||
/// program, the x8-equivalent operation count) in place of the mixer. Name "dr736". A prototype class; not v3.
|
||||
pub const DR736: LoadClass =
|
||||
LoadClass { mixer_mult: 1, derive_len: crate::derive::DERIVE_LEN_X8 as u16, ..LoadClass::MX4 };
|
||||
|
||||
/// This class with a derivation program of `len` instructions per round program (0: the fixed mixer); the
|
||||
/// mixer multiplier is set to 1 under a program, since no mixer is applied.
|
||||
pub fn with_derive(self, len: u16) -> LoadClass {
|
||||
assert!(len == 0 || (32..=4096).contains(&len), "derivation program length must be 0 or 32..=4096");
|
||||
LoadClass { derive_len: len, mixer_mult: if len != 0 { 1 } else { self.mixer_mult }, ..self }
|
||||
}
|
||||
|
||||
/// Whether the item derivation is the per-day program (Counter ASIC 3.0 item 2).
|
||||
pub fn is_derived(&self) -> bool {
|
||||
self.derive_len != 0
|
||||
}
|
||||
|
||||
/// A fixed width (1, 4 or 16 words) with `load_slots` loads per program.
|
||||
pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass {
|
||||
let mut mix = [0u8; 3];
|
||||
|
|
@ -505,6 +527,16 @@ impl LoadClass {
|
|||
if s == "mx8" {
|
||||
return Some(LoadClass::MX8);
|
||||
}
|
||||
// Counter ASIC 3.0 item 2: "dr<len>" is the derivation class on the v3 loads and growth rule
|
||||
if let Some(digits) = s.strip_prefix("dr") {
|
||||
if !digits.is_empty() && digits.bytes().all(|b| b.is_ascii_digit()) {
|
||||
let len: u16 = digits.parse().ok()?;
|
||||
if len == 0 || !(32..=4096).contains(&len) {
|
||||
return None;
|
||||
}
|
||||
return Some(LoadClass::MX4.with_derive(len));
|
||||
}
|
||||
}
|
||||
// the mixer suffix: "...m<mult>" then an optional "g"
|
||||
let (s, growth) = match s.strip_suffix('g') {
|
||||
Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true),
|
||||
|
|
@ -627,6 +659,11 @@ impl LoadClass {
|
|||
if *self == LoadClass::MX8 {
|
||||
return "mx8".to_string();
|
||||
}
|
||||
if self.derive_len != 0 {
|
||||
// the derivation class: "dr<len>" on the v3 loads; any other base keeps its name with the suffix
|
||||
let base = LoadClass { derive_len: 0, mixer_mult: 1, ..*self }.base_name();
|
||||
return if base == "v2m1g" || base == "v2" { format!("dr{}", self.derive_len) } else { format!("{base}dr{}", self.derive_len) };
|
||||
}
|
||||
let loads = LoadClass { mixer_mult: 1, growth: false, ..*self };
|
||||
let base = if loads.is_v2() {
|
||||
"v2".to_string()
|
||||
|
|
@ -892,6 +929,11 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L
|
|||
b.push(class.mixer_mult);
|
||||
b.push(class.growth as u8);
|
||||
}
|
||||
if class.derive_len != 0 {
|
||||
// Counter ASIC 3.0 item 2: the derivation program's length is part of the construction
|
||||
b.extend_from_slice(b"derive/");
|
||||
b.extend_from_slice(&class.derive_len.to_le_bytes());
|
||||
}
|
||||
if let Some(e) = class.era {
|
||||
b.extend_from_slice(b"era/");
|
||||
b.extend_from_slice(&e.id_bytes());
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@
|
|||
//! * [`accept`]: the acceptance rule every candidate program must pass; a rejected candidate is replaced by the
|
||||
//! next attempt of the same seed.
|
||||
//! * [`memhard`]: the 256 MiB ChaCha12 cache and the 8-round dataset item derivation (`proto-metal/MEMHARD.md`).
|
||||
//! * [`derive`]: the per-day item-derivation program of Counter ASIC 3.0 item 2 (a prototype behind a load class).
|
||||
//! * [`verify`]: the 32-lane warp interpreter that computes the 64-bit hash on the CPU, deriving dataset
|
||||
//! words on demand from the cache (or from the closed form, for the old packs).
|
||||
//! * [`emit`]: the Metal, CUDA and OpenCL kernel text for a program, byte-identical to the Swift exporter.
|
||||
|
|
@ -23,6 +24,7 @@
|
|||
|
||||
pub mod accept;
|
||||
pub mod bind;
|
||||
pub mod derive;
|
||||
pub mod emit;
|
||||
pub mod generator;
|
||||
pub mod memhard;
|
||||
|
|
|
|||
|
|
@ -93,7 +93,7 @@ fn usage() -> ! {
|
|||
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
|
||||
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
|
||||
\x20 show the accepted program, one instruction per line\n\
|
||||
\x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
|
||||
\x20 --class C load class: v2 (default), mx4, mx8 (class v3: mixer x8, cache growth), dr<len> (Counter ASIC 3.0 item 2: the per-day derivation program, dr736 = the x8-equivalent), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
|
||||
\x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)\n\
|
||||
\x20 --program-class v2|v3 the program class of the seam (v3 = generator 3 on V3_CLASS, the chain's own derivation; --era-hex records the era seed)\n\
|
||||
\x20 --era E era layout over --class: igneum-era-test/<n> or <n>:<64 hex> (the 32-byte era seed E_n)\n\
|
||||
|
|
@ -274,6 +274,20 @@ fn bench(a: &Args, mode: DatasetMode) {
|
|||
shape.cache_log2_words,
|
||||
e.program.op_mix()
|
||||
);
|
||||
if let Some(dp) = e.dataset.memhard().and_then(|m| m.params.derive.as_ref()) {
|
||||
// Counter ASIC 3.0 item 2: the day's derivation program
|
||||
println!(
|
||||
"derivation program: {} instructions per round program, {} per item ({} GPU ops, {} chip ops, {} multiplies per item), attempt {}, fingerprint {:016x}, op mix {}",
|
||||
dp.len,
|
||||
dp.instr_count(),
|
||||
dp.gpu_ops(),
|
||||
dp.chip_ops(),
|
||||
dp.muls(),
|
||||
dp.attempt,
|
||||
dp.fingerprint(),
|
||||
dp.op_mix()
|
||||
);
|
||||
}
|
||||
let bases = [0u32, 4096, 1_000_000];
|
||||
for &b in &bases {
|
||||
let t = Instant::now();
|
||||
|
|
|
|||
|
|
@ -9,6 +9,7 @@
|
|||
//! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the
|
||||
//! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit.
|
||||
|
||||
use crate::derive::{run_round, DeriveProgram, SoaState, DERIVE_REGS, SOA_LANES};
|
||||
use crate::generator::LoadClass;
|
||||
use crate::seed::{day_key, fnv1a64_words, SplitMix64};
|
||||
|
||||
|
|
@ -45,11 +46,15 @@ pub struct Shape {
|
|||
pub mixer_mult: u32,
|
||||
/// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3).
|
||||
pub cache_log2_words: u32,
|
||||
/// Counter ASIC 3.0 item 2 (`crate::derive`): instructions per round program of the per-day derivation
|
||||
/// program, which replaces the `mixer_mult` applications of `M_r` in every mixer slot when non-zero. 0 for
|
||||
/// version 2 and class v3 (the fixed mixer).
|
||||
pub derive_len: u32,
|
||||
}
|
||||
|
||||
impl Shape {
|
||||
/// Version 2: one mixer application per round, a 2^26-word cache.
|
||||
pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32 };
|
||||
pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32, derive_len: 0 };
|
||||
|
||||
/// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule).
|
||||
pub fn for_class(class: &LoadClass) -> Shape {
|
||||
|
|
@ -62,6 +67,7 @@ impl Shape {
|
|||
Shape {
|
||||
mixer_mult: class.mixer_mult(),
|
||||
cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 },
|
||||
derive_len: class.derive_len as u32,
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -83,9 +89,21 @@ impl Shape {
|
|||
pub fn log2_segments(&self) -> u32 {
|
||||
self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32
|
||||
}
|
||||
/// Mixer applications per item: `(ITEM_ROUNDS + 1) x m`.
|
||||
/// Mixer applications per item: `(ITEM_ROUNDS + 1) x m` (0 under a derivation program, which has no mixer).
|
||||
pub fn mixers_per_item(&self) -> u32 {
|
||||
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult
|
||||
if self.is_derived() {
|
||||
0
|
||||
} else {
|
||||
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult
|
||||
}
|
||||
}
|
||||
/// Whether the item derivation is the per-day program of `crate::derive` (Counter ASIC 3.0 item 2).
|
||||
pub fn is_derived(&self) -> bool {
|
||||
self.derive_len != 0
|
||||
}
|
||||
/// Instructions per item under a derivation program: `(ITEM_ROUNDS + 1) x derive_len`.
|
||||
pub fn derive_instrs_per_item(&self) -> u32 {
|
||||
(ITEM_ROUNDS as u32 + 1) * self.derive_len
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -169,7 +187,10 @@ pub fn chacha_block(x: &[u32; 16]) -> [u32; 16] {
|
|||
}
|
||||
|
||||
/// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order:
|
||||
/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's.
|
||||
/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's. Under a shape with
|
||||
/// a derivation program (Counter ASIC 3.0 item 2) the same stream continues after the 40 draws with the program's
|
||||
/// draws (`DeriveProgram::draw`); the mixer constants are still drawn (the item init uses MUL and RC) and the
|
||||
/// mixer itself is not applied.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct MixParams {
|
||||
pub key: [u32; 8],
|
||||
|
|
@ -177,6 +198,8 @@ pub struct MixParams {
|
|||
pub mul: [u32; 16],
|
||||
pub rc: [u32; 16],
|
||||
pub shape: Shape,
|
||||
/// The per-day derivation program when `shape.derive_len != 0`, else `None`.
|
||||
pub derive: Option<DeriveProgram>,
|
||||
}
|
||||
|
||||
impl MixParams {
|
||||
|
|
@ -198,7 +221,8 @@ impl MixParams {
|
|||
for c in rc.iter_mut() {
|
||||
*c = rng.next() as u32;
|
||||
}
|
||||
Self { key, rot, mul, rc, shape }
|
||||
let derive = if shape.is_derived() { Some(DeriveProgram::draw(&mut rng, shape.derive_len)) } else { None };
|
||||
Self { key, rot, mul, rc, shape, derive }
|
||||
}
|
||||
/// Parameters for a day string: the key is `seed_words("day/" + day)`.
|
||||
pub fn for_day(day: &str) -> Self {
|
||||
|
|
@ -350,7 +374,7 @@ impl Cache {
|
|||
/// are the smaller cache's segments word for word.
|
||||
pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache {
|
||||
assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30");
|
||||
let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words };
|
||||
let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words, derive_len: 0 };
|
||||
let mut words = vec![0u32; shape.cache_words()];
|
||||
for seg in 0..shape.cache_segments() {
|
||||
Self::fill_segment(&mut words, seg, &key);
|
||||
|
|
@ -481,6 +505,9 @@ impl HotTable {
|
|||
/// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its
|
||||
/// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2.
|
||||
pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) {
|
||||
if let Some(prog) = &mp.derive {
|
||||
return derive_items_program(ts, mp, prog, cache, out);
|
||||
}
|
||||
// The item loop lives in its own function, one instance per cache size the growth rule can reach with the line
|
||||
// mask a constant, never inlined into the callers. Inlined into `MemhardCpu::fetch` it ran at 1.33 ms per unit
|
||||
// against 0.61 out of line (the version 2 verifier, bisected on one core under the measure lock, 5 October 2026,
|
||||
|
|
@ -535,6 +562,44 @@ fn derive_items_mask<const LINE_MASK: u32>(ts: &[u32], mp: &MixParams, cache: &C
|
|||
}
|
||||
}
|
||||
|
||||
/// [`derive_items`] under a per-day derivation program (Counter ASIC 3.0 item 2, `crate::derive`): the same
|
||||
/// init and the same 8 dependent cache reads, with round program `r` in place of the `m` mixer applications of
|
||||
/// round `r` and program 8 in place of the final applications. The states are kept word-major
|
||||
/// (`st[reg][lane]`) so every instruction runs across the batch's items in one vectorised loop and the
|
||||
/// interpreter's dispatch is paid once per instruction per batch of up to 32 items, not once per item. The
|
||||
/// cache reads of the batch are issued together, as in the fixed-mixer loop, so the 8 dependent misses of
|
||||
/// independent items overlap in the memory system.
|
||||
#[inline(never)]
|
||||
pub fn derive_items_program(ts: &[u32], mp: &MixParams, prog: &DeriveProgram, cache: &Cache, out: &mut [[u32; 16]]) {
|
||||
let n = ts.len();
|
||||
debug_assert!(out.len() >= n && n <= SOA_LANES);
|
||||
assert_eq!(prog.rounds.len(), ITEM_ROUNDS + 1);
|
||||
let mut st: SoaState = [[0u32; SOA_LANES]; DERIVE_REGS];
|
||||
for k in 0..n {
|
||||
let t = ts[k];
|
||||
for i in 0..8 {
|
||||
st[i][k] = mp.key[i];
|
||||
st[8 + i][k] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
|
||||
}
|
||||
}
|
||||
let mask = cache.line_mask();
|
||||
for r in 0..ITEM_ROUNDS {
|
||||
run_round(&prog.rounds[r], &mut st);
|
||||
for k in 0..n {
|
||||
let line = cache.line(st[0][k] & mask);
|
||||
for i in 0..16 {
|
||||
st[i][k] ^= line[i];
|
||||
}
|
||||
}
|
||||
}
|
||||
run_round(&prog.rounds[ITEM_ROUNDS], &mut st);
|
||||
for k in 0..n {
|
||||
for i in 0..16 {
|
||||
out[k][i] = st[i][k];
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// One dataset item, 16 words.
|
||||
pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
|
||||
let mut out = [[0u32; 16]; 1];
|
||||
|
|
@ -786,7 +851,7 @@ mod tests {
|
|||
let v2 = Shape::for_class_day(&LoadClass::V2, 100_000);
|
||||
assert_eq!(v2, Shape::V2);
|
||||
let v3 = Shape::for_class_day(&LoadClass::MX4, 0);
|
||||
assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26 });
|
||||
assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26, derive_len: 0 });
|
||||
assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27);
|
||||
assert_eq!(v3.mixers_per_item(), 36);
|
||||
assert_eq!(Shape::V2.mixers_per_item(), 9);
|
||||
|
|
@ -807,7 +872,7 @@ mod tests {
|
|||
assert_eq!(small.segments(), 64);
|
||||
assert_eq!(small.line_mask(), 4095);
|
||||
for m in [1u32, 2, 4] {
|
||||
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16 });
|
||||
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16, derive_len: 0 });
|
||||
for t in [0u32, 1, 12_345, u32::MAX] {
|
||||
let got = derive_item(t, &mp, &small);
|
||||
let mut s = [0u32; 16];
|
||||
|
|
@ -830,8 +895,8 @@ mod tests {
|
|||
assert_eq!(got, s, "m {m} t {t}");
|
||||
}
|
||||
}
|
||||
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16 });
|
||||
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16 });
|
||||
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: 0 });
|
||||
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16, derive_len: 0 });
|
||||
assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small));
|
||||
assert_eq!(round_key_mult(0, 0, 1), round_key(0));
|
||||
assert_eq!(round_key_mult(8, 0, 1), round_key(8));
|
||||
|
|
|
|||
275
igneum-pow/tests/derive.rs
Normal file
275
igneum-pow/tests/derive.rs
Normal file
|
|
@ -0,0 +1,275 @@
|
|||
//! The per-day item-derivation program (Counter ASIC 3.0 item 2, `docs/plans/counter-asic-3-derivation.md`): the
|
||||
//! soundness runs on the CPU.
|
||||
//!
|
||||
//! 1. By hand: the derived item of class dr736 restated with the scalar reference on a small cache equals the
|
||||
//! verifier's batched (SoA) derivation, at the item boundary and the index wrap.
|
||||
//! 2. The v2 and v3 paths are untouched: the derivation field is 0 on both, their items are the fixed-mixer items.
|
||||
//! 3. Determinism and the stream: two epochs agree on every vector and file; the program is the continuation of the
|
||||
//! mixer's stream (the first draw of the program follows the 40 mixer draws); a different day draws a different
|
||||
//! program; the class is in the program id and the name.
|
||||
//! 4. Stats beside x8: bit balance and single-bit avalanche of the derived items and of the hash, on the same seeds.
|
||||
//! 5. A program whose text is the kernels': every instruction's C text evaluated by hand on one state matches the
|
||||
//! scalar reference (the text forms are what Metal, CUDA and OpenCL compile).
|
||||
|
||||
use igneum_pow::derive::{instr_text, run_round_scalar, DOp, DeriveProgram, DERIVE_LEN_X8, DERIVE_PROGRAMS};
|
||||
use igneum_pow::emit::export_pack;
|
||||
use igneum_pow::generator::{generate_from_seed_bytes_class, LoadClass, V3_CLASS};
|
||||
use igneum_pow::memhard::{derive_item, derive_items, mixer, round_key, Cache, MixParams, Shape, ITEM_ROUNDS};
|
||||
use igneum_pow::seed::{day_key, SplitMix64};
|
||||
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
|
||||
|
||||
const DAY: &str = "2026-10-03";
|
||||
|
||||
/// The item of a derivation class restated by hand with the scalar reference.
|
||||
fn item_by_hand(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
|
||||
let prog = mp.derive.as_ref().expect("a derivation class");
|
||||
let mut s = [0u32; 16];
|
||||
s[..8].copy_from_slice(&mp.key);
|
||||
for i in 0..8 {
|
||||
s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
|
||||
}
|
||||
for r in 0..ITEM_ROUNDS {
|
||||
run_round_scalar(&prog.rounds[r], &mut s);
|
||||
let line = cache.line(s[0]);
|
||||
for i in 0..16 {
|
||||
s[i] ^= line[i];
|
||||
}
|
||||
}
|
||||
run_round_scalar(&prog.rounds[ITEM_ROUNDS], &mut s);
|
||||
s
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn derived_item_by_hand_and_in_batches() {
|
||||
let key = day_key(DAY);
|
||||
let cache = Cache::fill_log2(key, 16);
|
||||
let shape = Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: DERIVE_LEN_X8 };
|
||||
let mp = MixParams::with_shape(key, shape);
|
||||
let prog = mp.derive.as_ref().unwrap();
|
||||
assert_eq!(prog.rounds.len(), DERIVE_PROGRAMS);
|
||||
assert!(prog.check().is_ok());
|
||||
assert_eq!(shape.derive_instrs_per_item(), 9 * DERIVE_LEN_X8);
|
||||
assert_eq!(shape.mixers_per_item(), 0);
|
||||
for t in [0u32, 1, 2, 15, 16, 17, 12_345, (1 << 28) - 1, u32::MAX - 1, u32::MAX] {
|
||||
assert_eq!(derive_item(t, &mp, &cache), item_by_hand(t, &mp, &cache), "t {t}");
|
||||
}
|
||||
// a batch of 32 distinct items against one at a time, and a short batch
|
||||
let ts: Vec<u32> = (0..32).map(|k| k * 7_919 + 3).collect();
|
||||
let mut out = [[0u32; 16]; 32];
|
||||
derive_items(&ts, &mp, &cache, &mut out);
|
||||
for (k, &t) in ts.iter().enumerate() {
|
||||
assert_eq!(out[k], item_by_hand(t, &mp, &cache), "batch slot {k}");
|
||||
}
|
||||
let mut out5 = [[0u32; 16]; 5];
|
||||
derive_items(&ts[..5], &mp, &cache, &mut out5);
|
||||
assert_eq!(&out5[..], &out[..5]);
|
||||
// the fixed mixer of the same key gives other items
|
||||
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 8, cache_log2_words: 16, derive_len: 0 });
|
||||
assert!(v3.derive.is_none());
|
||||
assert_ne!(derive_item(0, &v3, &cache), derive_item(0, &mp, &cache));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn v2_and_v3_are_untouched() {
|
||||
assert_eq!(LoadClass::V2.derive_len, 0);
|
||||
assert_eq!(V3_CLASS.derive_len, 0);
|
||||
assert_eq!(LoadClass::MX8.derive_len, 0);
|
||||
assert_eq!(Shape::V2.derive_len, 0);
|
||||
assert!(!Shape::for_class(&V3_CLASS).is_derived());
|
||||
let key = day_key(DAY);
|
||||
let cache = Cache::fill_log2(key, 16);
|
||||
// the version 2 item restated by hand (the mixer_mult_by_hand test of memhard.rs, m = 1)
|
||||
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: 0 });
|
||||
let t = 12_345u32;
|
||||
let mut s = [0u32; 16];
|
||||
s[..8].copy_from_slice(&key);
|
||||
for i in 0..8 {
|
||||
s[8 + i] = t.wrapping_mul(v2.mul[i]).wrapping_add(v2.rc[i]);
|
||||
}
|
||||
for r in 0..8usize {
|
||||
mixer(&mut s, round_key(r), &v2);
|
||||
let line = cache.line(s[0]);
|
||||
for i in 0..16 {
|
||||
s[i] ^= line[i];
|
||||
}
|
||||
}
|
||||
mixer(&mut s, round_key(8), &v2);
|
||||
assert_eq!(derive_item(t, &v2, &cache), s);
|
||||
// the mixer constants of the derivation class are the v2 draws (the stream continues after them)
|
||||
let dr = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: DERIVE_LEN_X8 });
|
||||
assert_eq!((dr.rot, dr.mul, dr.rc), (v2.rot, v2.mul, v2.rc));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stream_class_name_and_id() {
|
||||
let key = day_key(DAY);
|
||||
// the program is the continuation of the mixer stream: 40 draws, then the program
|
||||
let mut rng = SplitMix64::new(key[0] as u64 | ((key[1] as u64) << 32));
|
||||
for _ in 0..40 {
|
||||
rng.next();
|
||||
}
|
||||
let expect = DeriveProgram::draw(&mut rng, DERIVE_LEN_X8);
|
||||
let mp = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 26, derive_len: DERIVE_LEN_X8 });
|
||||
assert_eq!(mp.derive.as_ref().unwrap(), &expect);
|
||||
// another day, another program; another length, another program
|
||||
let other = MixParams::with_shape(day_key("2026-10-04"), Shape { mixer_mult: 1, cache_log2_words: 26, derive_len: DERIVE_LEN_X8 });
|
||||
assert_ne!(other.derive.as_ref().unwrap().fingerprint(), expect.fingerprint());
|
||||
let short = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 26, derive_len: 368 });
|
||||
assert_eq!(short.derive.as_ref().unwrap().instr_count(), 9 * 368);
|
||||
// the class: name, parse, id, and the v2 program stream (v2 loads, no width roll)
|
||||
let c = LoadClass::DR736;
|
||||
assert_eq!(c.name(), "dr736");
|
||||
assert_eq!(LoadClass::parse("dr736"), Some(c));
|
||||
assert_eq!(LoadClass::parse("dr368"), Some(LoadClass::MX4.with_derive(368)));
|
||||
assert_eq!(LoadClass::parse("dr0"), None);
|
||||
assert_eq!(LoadClass::parse("dr5000"), None);
|
||||
assert!(c.v2_loads() && !c.takes_width_roll() && c.growth && c.mixer_mult == 1 && c.is_derived());
|
||||
let p = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", c);
|
||||
let v2 = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", LoadClass::V2);
|
||||
let mx8 = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", LoadClass::MX8);
|
||||
assert_eq!(p.instrs, v2.instrs, "the v2 program of the seed");
|
||||
assert_ne!(p.program_id(), v2.program_id());
|
||||
assert_ne!(p.program_id(), mx8.program_id());
|
||||
assert_ne!(
|
||||
generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", LoadClass::MX4.with_derive(368)).program_id(),
|
||||
p.program_id()
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn determinism_and_pack_text() {
|
||||
let a = Epoch::new_class("igneum-genesis", DAY, DatasetMode::MemoryHard, 20, LoadClass::DR736);
|
||||
let b = Epoch::new_class("igneum-genesis", DAY, DatasetMode::MemoryHard, 20, LoadClass::DR736);
|
||||
let pa = export_pack(&a, DAY, "test");
|
||||
let pb = export_pack(&b, DAY, "test");
|
||||
assert_eq!(pa.outs, pb.outs);
|
||||
assert_eq!(pa.files, pb.files);
|
||||
let names: Vec<&str> = pa.files.iter().map(|(n, _)| n.as_str()).collect();
|
||||
assert!(names.contains(&"memhard.h") && names.contains(&"memhard.metal"));
|
||||
for (name, text) in &pa.files {
|
||||
if name == "memhard.h" || name == "memhard.metal" || name == "kernel.cl" {
|
||||
for r in 0..DERIVE_PROGRAMS {
|
||||
assert!(text.contains(&format!("mh_round_{r}(")), "{name} carries round program {r}");
|
||||
}
|
||||
assert!(!text.contains("mh_mixer(s,"), "{name}: no mixer application under a derivation program");
|
||||
let dp = a.dataset.memhard().unwrap().params.derive.as_ref().unwrap();
|
||||
// every instruction's text appears, in order, inside the round functions
|
||||
let first = instr_text(&dp.rounds[0][0]);
|
||||
assert!(text.contains(&first), "{name} carries the first instruction {first}");
|
||||
}
|
||||
if name == "program.h" {
|
||||
assert!(text.contains("#define IGNEUM_LOAD_CLASS \"dr736\""));
|
||||
assert!(text.contains("#define IGNEUM_DERIVE_LEN 736"));
|
||||
assert!(text.contains("#define IGNEUM_CLASS_DERIVE_LEN 736"));
|
||||
assert!(!text.contains("#define IGNEUM_MIXER_MULT"));
|
||||
assert!(text.contains("#define IGNEUM_GENERATOR 2"));
|
||||
}
|
||||
if name == "program.json" {
|
||||
assert!(text.contains("\"derive_len\": 736"));
|
||||
assert!(text.contains("\"programs\": ["));
|
||||
serde_json::from_str::<serde_json::Value>(text).expect("valid JSON");
|
||||
}
|
||||
}
|
||||
// the vectors are the interpreter's
|
||||
assert_eq!(pa.outs[0], a.hash_warp(0));
|
||||
}
|
||||
|
||||
/// Bit balance and single-bit avalanche of the derived items (t flipped one bit) and of the hash, beside x8.
|
||||
#[test]
|
||||
fn stats_beside_x8() {
|
||||
let key = day_key(DAY);
|
||||
let cache = Cache::fill_log2(key, 18);
|
||||
let dr = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 18, derive_len: DERIVE_LEN_X8 });
|
||||
let x8 = MixParams::with_shape(key, Shape { mixer_mult: 8, cache_log2_words: 18, derive_len: 0 });
|
||||
for (label, mp) in [("dr736", &dr), ("x8", &x8)] {
|
||||
let n = 2048u32;
|
||||
let mut ones = [0u32; 512];
|
||||
let mut flips = 0u64;
|
||||
let mut flip_n = 0u64;
|
||||
let mut seen = std::collections::HashSet::new();
|
||||
for t in 0..n {
|
||||
let a = derive_item(t * 2_654_435_761, mp, &cache);
|
||||
assert!(seen.insert(a), "{label}: duplicate item");
|
||||
for i in 0..16 {
|
||||
for b in 0..32 {
|
||||
ones[i * 32 + b] += (a[i] >> b) & 1;
|
||||
}
|
||||
}
|
||||
let bit = t % 32;
|
||||
let c = derive_item((t * 2_654_435_761) ^ (1 << bit), mp, &cache);
|
||||
for i in 0..16 {
|
||||
flips += (a[i] ^ c[i]).count_ones() as u64;
|
||||
}
|
||||
flip_n += 512;
|
||||
}
|
||||
let avalanche = flips as f64 / flip_n as f64 * 100.0;
|
||||
let worst_z = ones.iter().map(|&o| ((o as f64 - n as f64 / 2.0) / (n as f64 / 4.0).sqrt()).abs()).fold(0.0, f64::max);
|
||||
println!("{label}: avalanche {avalanche:.2} percent, worst bit z {worst_z:.2}");
|
||||
assert!((48.5..=51.5).contains(&avalanche), "{label}: avalanche {avalanche}");
|
||||
assert!(worst_z < 4.5, "{label}: worst bit z {worst_z}");
|
||||
}
|
||||
// the hash on the same seeds: avalanche across a nonce flip within the unit
|
||||
let e = Epoch::new_class("igneum-genesis", DAY, DatasetMode::MemoryHard, 20, LoadClass::DR736);
|
||||
let w0 = e.hash_warp(0);
|
||||
let w1 = e.hash_warp(32);
|
||||
let mut d = 0u64;
|
||||
for l in 0..32 {
|
||||
d += (w0[l] ^ w1[l]).count_ones() as u64;
|
||||
}
|
||||
let av = d as f64 / (32.0 * 64.0) * 100.0;
|
||||
println!("hash unit 0 against unit 1: {av:.2} percent of bits differ");
|
||||
assert!((44.0..=56.0).contains(&av));
|
||||
}
|
||||
|
||||
/// The kernels' text forms: every form's C text, read back into the scalar reference's arithmetic by hand.
|
||||
#[test]
|
||||
fn text_forms_match_scalar_reference() {
|
||||
let mut rng = SplitMix64::new(42);
|
||||
let p = DeriveProgram::draw_candidate(&mut rng, 64, 0);
|
||||
let mut s = [0u32; 16];
|
||||
for (i, v) in s.iter_mut().enumerate() {
|
||||
*v = 0x9e37_79b9u32.wrapping_mul(i as u32 + 1);
|
||||
}
|
||||
let mut seen = [false; 12];
|
||||
for ins in p.rounds.iter().flatten() {
|
||||
seen[ins.op as usize] = true;
|
||||
let before = s;
|
||||
run_round_scalar(std::slice::from_ref(ins), &mut s);
|
||||
let (d, c, b) = (ins.dst as usize, ins.src as usize, ins.src2 as usize);
|
||||
let (dv, cv, bv) = (before[d], before[c], before[b]);
|
||||
let want = match ins.op {
|
||||
DOp::Add => dv.wrapping_add(cv),
|
||||
DOp::Sub => dv.wrapping_sub(cv),
|
||||
DOp::Xor => dv ^ cv,
|
||||
DOp::Mul => dv.wrapping_mul(cv | 1),
|
||||
DOp::Rot => dv.rotate_left(ins.rot as u32).wrapping_add(cv),
|
||||
DOp::XRot => (dv ^ cv).rotate_left(ins.rot as u32),
|
||||
DOp::AddC => dv.wrapping_add(cv.wrapping_add(ins.imm)),
|
||||
DOp::XorC => dv ^ cv ^ ins.imm,
|
||||
DOp::MulC => (dv ^ cv).wrapping_mul(ins.imm),
|
||||
DOp::MulC2 => dv.wrapping_mul(ins.imm).wrapping_add(cv),
|
||||
DOp::AndX => dv ^ (cv & bv),
|
||||
DOp::OrX => dv.wrapping_add(cv | bv),
|
||||
};
|
||||
assert_eq!(s[d], want, "{}", instr_text(ins));
|
||||
for i in 0..16 {
|
||||
if i != d {
|
||||
assert_eq!(s[i], before[i], "only the destination changes: {}", instr_text(ins));
|
||||
}
|
||||
}
|
||||
assert!(instr_text(ins).starts_with(&format!("s[{d}]")));
|
||||
}
|
||||
assert!(seen.iter().all(|s| *s), "64 x 9 draws cover every form");
|
||||
}
|
||||
|
||||
/// The dataset source of the class on a day: the verifier's `word` path derives through the program.
|
||||
#[test]
|
||||
fn dataset_source_word_path() {
|
||||
let ds = DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 20, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: DERIVE_LEN_X8 });
|
||||
let m = ds.memhard().unwrap();
|
||||
let item = derive_item(3, &m.params, &m.cache);
|
||||
for j in 0..16u32 {
|
||||
assert_eq!(ds.word(3 * 16 + j), item[j as usize]);
|
||||
}
|
||||
assert_eq!(ds.word((1 << 20) + 5), ds.word(5), "the mask");
|
||||
}
|
||||
|
|
@ -181,7 +181,7 @@ fn edge_items_every_multiplier() {
|
|||
let key = day_key(DAY);
|
||||
let cache = Cache::fill_log2(key, 14);
|
||||
for m in [1u32, 2, 4, 8] {
|
||||
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 14 });
|
||||
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 14, derive_len: 0 });
|
||||
let by_hand = |t: u32| -> [u32; 16] {
|
||||
let mut s = [0u32; 16];
|
||||
s[..8].copy_from_slice(&key);
|
||||
|
|
|
|||
|
|
@ -369,7 +369,7 @@ fn v3_packs_are_the_v2_seeds_under_mixer_x8() {
|
|||
assert_eq!(e3.program.program_id(), igneum_pow::generator::program_id(GENERATOR_VERSION_V3, &e3.program.seed, e3.program.attempt));
|
||||
let m3 = e3.dataset.memhard().unwrap();
|
||||
let m2 = e2.dataset.memhard().unwrap();
|
||||
assert_eq!(m3.shape(), Shape { mixer_mult: 8, cache_log2_words: 26 });
|
||||
assert_eq!(m3.shape(), Shape { mixer_mult: 8, cache_log2_words: 26, derive_len: 0 });
|
||||
assert_eq!(m3.cache.fnv1a64(), m2.cache.fnv1a64(), "{v3}: the same cache as v2 on day 0");
|
||||
assert_eq!(m3.params.rot, m2.params.rot);
|
||||
assert_eq!(e3.dataset.log2_words, 28);
|
||||
|
|
@ -405,7 +405,7 @@ fn v3_packs_are_the_v2_seeds_under_mixer_x8() {
|
|||
assert_eq!(j["load_class"].as_str().unwrap(), "mx4");
|
||||
assert_eq!(e4.program.class, LoadClass::MX4);
|
||||
assert_eq!(e4.program.instrs, epoch(v2).program.instrs);
|
||||
assert_eq!(e4.dataset.memhard().unwrap().shape(), Shape { mixer_mult: 4, cache_log2_words: 26 });
|
||||
assert_eq!(e4.dataset.memhard().unwrap().shape(), Shape { mixer_mult: 4, cache_log2_words: 26, derive_len: 0 });
|
||||
assert!(read(x4, "memhard.h").contains("j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 4u + j + 1u))"));
|
||||
}
|
||||
}
|
||||
|
|
|
|||
6943
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel.cl
Normal file
6943
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel.cl
Normal file
File diff suppressed because it is too large
Load diff
164
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel.cu
Normal file
164
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = __umulhi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = __umulhi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = __umulhi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = __umulhi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = __umulhi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = __umulhi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = __umulhi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = __umulhi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = __umulhi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
7037
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel_bound.cl
Normal file
7037
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel_bound.cl
Normal file
File diff suppressed because it is too large
Load diff
123
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = __umulhi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = __umulhi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = __umulhi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = __umulhi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = __umulhi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = __umulhi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = __umulhi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = __umulhi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = __umulhi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
6773
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/memhard.h
Normal file
6773
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/memhard.h
Normal file
File diff suppressed because it is too large
Load diff
6771
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/memhard.metal
Normal file
6771
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/memhard.metal
Normal file
File diff suppressed because it is too large
Load diff
72
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/program.h
Normal file
72
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/program.h
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 2
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x7f4a5ca0a3637820ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=13 mulhi=9 shfl=5 sub=5 mad=4 rotl=4 xor=3 or=2 rotr=2 mul=1"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "dr736"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 1
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
// Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md, a prototype, NOT class v3): the item
|
||||
// derivation runs the day's drawn program (memhard.h: mh_round_0..8, IGNEUM_DERIVE_LEN instructions each) in place of the mixer.
|
||||
#define IGNEUM_CLASS_DERIVE_LEN 736
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_DERIVE_LEN 736 // instructions per round program of the day's item-derivation program (Counter ASIC 3.0 item 2; memhard.h mh_round_0..8)
|
||||
#define IGNEUM_DERIVE_ATTEMPT 0
|
||||
#define IGNEUM_DERIVE_FINGERPRINT 0x771868df4e64d6abull
|
||||
#define IGNEUM_DERIVE_INSTRS_PER_ITEM 6624
|
||||
#define IGNEUM_DERIVE_GPU_OPS_PER_ITEM 10708
|
||||
#define IGNEUM_DERIVE_CHIP_OPS_PER_ITEM 10072
|
||||
#define IGNEUM_DERIVE_MULS_PER_ITEM 1475
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
151
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/program.json
Normal file
151
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/program.json
Normal file
File diff suppressed because one or more lines are too long
109
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/program.metal
Normal file
109
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
|
||||
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
|
||||
r0 = rotl_imm(r0, 19u); // 3
|
||||
r7 = rotr_var(r7, r6); // 4
|
||||
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
|
||||
r1 = mulhi(r1, r7); // 6
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 7
|
||||
r7 = r7 ^ dataset[r4 & MASK]; // 8
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 9
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 10
|
||||
r1 = r1 ^ dataset[r5 & MASK]; // 11
|
||||
r3 = mulhi(r3, r5); // 12
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 13
|
||||
r0 = r0 - r3; // 14
|
||||
r5 = r1 * r3 + r5; // 15
|
||||
r6 = mulhi(r6, r1); // 16
|
||||
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
|
||||
r0 = mulhi(r0, r6); // 18
|
||||
r5 = rotr_var(r5, r3); // 19
|
||||
r5 = mulhi(r5, r2); // 20
|
||||
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
|
||||
r1 = mulhi(r1, r5); // 23
|
||||
r2 = r2 - r5; // 24
|
||||
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
|
||||
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
|
||||
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r2 = r2 ^ dataset[r5 & MASK]; // 31
|
||||
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
|
||||
r4 = r5 * r7 + r4; // 33
|
||||
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
|
||||
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
|
||||
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
|
||||
r5 = r5 ^ r7; // 37
|
||||
r2 = r2 | r1; // 38
|
||||
r1 = mulhi(r1, r0); // 39
|
||||
r6 = rotl_imm(r6, 19u); // 40
|
||||
r4 = mulhi(r4, r6); // 41
|
||||
r6 = r6 - r0; // 42
|
||||
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 44
|
||||
r1 = r1 ^ r3; // 45
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 46
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 47
|
||||
r5 = r5 * r3; // 48
|
||||
r1 = r1 - r5; // 49
|
||||
r2 = rotl_imm(r2, 8u); // 50
|
||||
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 52
|
||||
r2 = r2 - r7; // 53
|
||||
r4 = r4 ^ r0; // 54
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 56
|
||||
r0 = r1 * r4 + r0; // 57
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 58
|
||||
r5 = r5 | r6; // 59
|
||||
r6 = r5 * r7 + r6; // 60
|
||||
r4 = rotl_imm(r4, 28u); // 61
|
||||
r5 = mulhi(r5, r0); // 62
|
||||
r3 = r3 ^ dataset[r6 & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
|
||||
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
|
||||
r0 = rotl_imm(r0, 19u); // 3
|
||||
r7 = rotr_var(r7, r6); // 4
|
||||
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
|
||||
r1 = mulhi(r1, r7); // 6
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 7
|
||||
r7 = r7 ^ dataset[r4 & MASK]; // 8
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 9
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 10
|
||||
r1 = r1 ^ dataset[r5 & MASK]; // 11
|
||||
r3 = mulhi(r3, r5); // 12
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 13
|
||||
r0 = r0 - r3; // 14
|
||||
r5 = r1 * r3 + r5; // 15
|
||||
r6 = mulhi(r6, r1); // 16
|
||||
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
|
||||
r0 = mulhi(r0, r6); // 18
|
||||
r5 = rotr_var(r5, r3); // 19
|
||||
r5 = mulhi(r5, r2); // 20
|
||||
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
|
||||
r1 = mulhi(r1, r5); // 23
|
||||
r2 = r2 - r5; // 24
|
||||
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
|
||||
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
|
||||
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r2 = r2 ^ dataset[r5 & MASK]; // 31
|
||||
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
|
||||
r4 = r5 * r7 + r4; // 33
|
||||
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
|
||||
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
|
||||
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
|
||||
r5 = r5 ^ r7; // 37
|
||||
r2 = r2 | r1; // 38
|
||||
r1 = mulhi(r1, r0); // 39
|
||||
r6 = rotl_imm(r6, 19u); // 40
|
||||
r4 = mulhi(r4, r6); // 41
|
||||
r6 = r6 - r0; // 42
|
||||
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 44
|
||||
r1 = r1 ^ r3; // 45
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 46
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 47
|
||||
r5 = r5 * r3; // 48
|
||||
r1 = r1 - r5; // 49
|
||||
r2 = rotl_imm(r2, 8u); // 50
|
||||
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 52
|
||||
r2 = r2 - r7; // 53
|
||||
r4 = r4 ^ r0; // 54
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 56
|
||||
r0 = r1 * r4 + r0; // 57
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 58
|
||||
r5 = r5 | r6; // 59
|
||||
r6 = r5 * r7 + r6; // 60
|
||||
r4 = rotl_imm(r4, 28u); // 61
|
||||
r5 = mulhi(r5, r0); // 62
|
||||
r3 = r3 ^ dataset[r6 & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
57
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/vectors.h
Normal file
57
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x435cd45db22fc83bull, 0xeb08b0eec4635b92ull, 0x5dac8870aed5c678ull, 0x2779f39382882c13ull, 0x465e11a6d64c87f4ull, 0x1a379f4f699ca9aeull, 0x03c6a8eb7ff849ebull, 0xbc1b878d6c277748ull,
|
||||
0x69adeb1f2d29c0ebull, 0x375894dbc9868b50ull, 0x79d2d8527b76b932ull, 0x4dcefb4f0a2ec3fcull, 0xe5ad9fb702dd5966ull, 0xa296d3e34edcfbc6ull, 0xc5c839c600924f03ull, 0xdc7228f16532bfbeull,
|
||||
0x81f595f907c9cb04ull, 0x1152850f76d4e90aull, 0x25446f27b34b11a9ull, 0xee986312f9e4e522ull, 0x3c5baa9645d0c99full, 0x144aff659162feb9ull, 0x497c9200b23e8e07ull, 0xc92501e7f6157dfeull,
|
||||
0x8caf96daa4ca385aull, 0x00062ab96f31e514ull, 0x9508e98517444c91ull, 0x8a3f948cd7fb9797ull, 0xab156a88a1b77e41ull, 0x2408a16a50352fa9ull, 0x2237f2014a735ac3ull, 0x7da1a30d3b416f06ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0xb868dcaeacd8e012ull, 0xe88143fdca4bfab2ull, 0x9ee48b8d044a1dbeull, 0xc62e9314832773f3ull, 0x998beb89868141d7ull, 0xb730b261ff57cc38ull, 0xb886beb356bc26edull, 0x7472e83c4b9f208cull,
|
||||
0x3c436e168b67d002ull, 0xb3c4c537b341c70full, 0xeea62a85a5c4d571ull, 0x31fb743150950537ull, 0x23d68caf4fb30f60ull, 0x7fa71a956c2cdab0ull, 0x2aaac765ae72764full, 0x5f109809ed36b7fbull,
|
||||
0xebde9ca62f1a205cull, 0xa684f0ab53357506ull, 0x5b68d0c5246705a2ull, 0x923549e6f6aad924ull, 0x056ad95d6c1b0cffull, 0x96879d74b28d43ddull, 0xc7cb03983950f903ull, 0x09391d92779b4915ull,
|
||||
0x92c061ed4bf9de47ull, 0xee84d54a38b36d07ull, 0x78da496aa8c9afd7ull, 0x0afe22f66ce1fc6eull, 0xf9576b3c7dc747c2ull, 0x16bfb648d41c69a2ull, 0xbf8f0934f1b1c327ull, 0x3781f5d303ba67f8ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0xcd1ed6c453bf0f8cull, 0x19a6d570facc3c28ull, 0x99a0491a592d2705ull, 0x84fe73aa729eeaebull, 0x05d6496a724a48bcull, 0x427d4a53e3cd224bull, 0x566c3ffcc3dcfd37ull, 0x8a7edbb1da3510c0ull,
|
||||
0x845f1b6d3c02f911ull, 0x9dc5b77664c6fb53ull, 0x7dc5df7872141de5ull, 0x281dd1300d03f493ull, 0xe56337a62077fb83ull, 0xbf3f0ceccb20a1e3ull, 0x61dd42e89af96505ull, 0x012b28dac792c2ffull,
|
||||
0x703e2d57a7228db9ull, 0x93617feef48f6fa0ull, 0xa9b76a0cc401f839ull, 0xdb06176d832ccce9ull, 0xf1c45184eef7624bull, 0x2b6dc1cd2eae9856ull, 0xebb01b8100e4b19cull, 0x7f9305e00616dfe9ull,
|
||||
0xdccda8f0b817fc71ull, 0x3a927aa1fd8d9bc8ull, 0xa4d82662a5fc6d5bull, 0x126c846254679b8bull, 0x8d792378f29a8a2bull, 0xa79da48e070f183full, 0xdfc380765068d276ull, 0x84aee1d7f9c8b2a6ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x0da55599u, 0x16ebf912u, 0x26ebba50u, 0x644a8670u, 0xb86da339u, 0xfdea8ce8u, 0xd8d38a59u, 0x5f930123u,
|
||||
0xc99284a8u, 0xfa9fa963u, 0xfd990357u, 0x8d1aad8cu, 0xc5c03c95u, 0x82af7120u, 0x9fbf6800u, 0x913cc5e8u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0x7c357826u;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0xf4b0010au, 0xc82a3ca4u, 0xfc72371bu, 0x4c31b63cu, 0x0f0d86b4u, 0x0860e638u, 0x6016ffc1u, 0xf1cac2c4u, 0xb520efedu, 0x9eac76d6u, 0xca4ccfe1u, 0xa07e7108u, 0xc1c60eacu, 0xfdd31f2du, 0x2225e701u, 0x71f7fde6u, 0xb52aed34u, 0x75a3dc19u, 0xfe023726u, 0x308b9e1du, 0x6003222cu, 0xa5cc70a2u, 0x22c2940fu, 0xdbcd5db2u, 0xeed44d67u, 0xfca48f4cu, 0x2744926eu, 0x24f03506u, 0xa90cad7du, 0x47db755eu, 0x2051b7a9u, 0x18365d2du, 0x69c0d7e2u, 0x0f1f31c6u, 0x912e45bfu, 0x03e4035eu, 0xa30eb890u, 0xf65796a0u, 0x0b97d1e6u, 0xb0014ea8u, 0xa02bb1dfu, 0x87476bf2u, 0xc2f22b36u, 0x2f0452a5u, 0xb6d0c4c3u, 0xc6a765e2u, 0xf2b8cbd5u, 0x928e50adu, 0x85dbe732u, 0x2b9ee10cu, 0x2a1c9124u, 0x7f9a88e2u, 0x92da37eau, 0x519214cbu, 0x7869844eu, 0x19062ee1u, 0x045f0735u, 0x09f5e114u, 0x6082c8d4u, 0x6bc68181u, 0x700987d5u, 0x3bca401du, 0x7829f7aau, 0xe325295cu
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/vectors.json
Normal file
36
proto-cuda/packs-ca3-derive/dr736-devnet-epoch0/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x435cd45db22fc83b", "0xeb08b0eec4635b92", "0x5dac8870aed5c678", "0x2779f39382882c13", "0x465e11a6d64c87f4", "0x1a379f4f699ca9ae", "0x03c6a8eb7ff849eb", "0xbc1b878d6c277748",
|
||||
"0x69adeb1f2d29c0eb", "0x375894dbc9868b50", "0x79d2d8527b76b932", "0x4dcefb4f0a2ec3fc", "0xe5ad9fb702dd5966", "0xa296d3e34edcfbc6", "0xc5c839c600924f03", "0xdc7228f16532bfbe",
|
||||
"0x81f595f907c9cb04", "0x1152850f76d4e90a", "0x25446f27b34b11a9", "0xee986312f9e4e522", "0x3c5baa9645d0c99f", "0x144aff659162feb9", "0x497c9200b23e8e07", "0xc92501e7f6157dfe",
|
||||
"0x8caf96daa4ca385a", "0x00062ab96f31e514", "0x9508e98517444c91", "0x8a3f948cd7fb9797", "0xab156a88a1b77e41", "0x2408a16a50352fa9", "0x2237f2014a735ac3", "0x7da1a30d3b416f06"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0xb868dcaeacd8e012", "0xe88143fdca4bfab2", "0x9ee48b8d044a1dbe", "0xc62e9314832773f3", "0x998beb89868141d7", "0xb730b261ff57cc38", "0xb886beb356bc26ed", "0x7472e83c4b9f208c",
|
||||
"0x3c436e168b67d002", "0xb3c4c537b341c70f", "0xeea62a85a5c4d571", "0x31fb743150950537", "0x23d68caf4fb30f60", "0x7fa71a956c2cdab0", "0x2aaac765ae72764f", "0x5f109809ed36b7fb",
|
||||
"0xebde9ca62f1a205c", "0xa684f0ab53357506", "0x5b68d0c5246705a2", "0x923549e6f6aad924", "0x056ad95d6c1b0cff", "0x96879d74b28d43dd", "0xc7cb03983950f903", "0x09391d92779b4915",
|
||||
"0x92c061ed4bf9de47", "0xee84d54a38b36d07", "0x78da496aa8c9afd7", "0x0afe22f66ce1fc6e", "0xf9576b3c7dc747c2", "0x16bfb648d41c69a2", "0xbf8f0934f1b1c327", "0x3781f5d303ba67f8"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0xcd1ed6c453bf0f8c", "0x19a6d570facc3c28", "0x99a0491a592d2705", "0x84fe73aa729eeaeb", "0x05d6496a724a48bc", "0x427d4a53e3cd224b", "0x566c3ffcc3dcfd37", "0x8a7edbb1da3510c0",
|
||||
"0x845f1b6d3c02f911", "0x9dc5b77664c6fb53", "0x7dc5df7872141de5", "0x281dd1300d03f493", "0xe56337a62077fb83", "0xbf3f0ceccb20a1e3", "0x61dd42e89af96505", "0x012b28dac792c2ff",
|
||||
"0x703e2d57a7228db9", "0x93617feef48f6fa0", "0xa9b76a0cc401f839", "0xdb06176d832ccce9", "0xf1c45184eef7624b", "0x2b6dc1cd2eae9856", "0xebb01b8100e4b19c", "0x7f9305e00616dfe9",
|
||||
"0xdccda8f0b817fc71", "0x3a927aa1fd8d9bc8", "0xa4d82662a5fc6d5b", "0x126c846254679b8b", "0x8d792378f29a8a2b", "0xa79da48e070f183f", "0xdfc380765068d276", "0x84aee1d7f9c8b2a6"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x0da55599", "0x16ebf912", "0x26ebba50", "0x644a8670", "0xb86da339", "0xfdea8ce8", "0xd8d38a59", "0x5f930123", "0xc99284a8", "0xfa9fa963", "0xfd990357", "0x8d1aad8c", "0xc5c03c95", "0x82af7120", "0x9fbf6800", "0x913cc5e8"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0x7c357826",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0xf4b0010a"}, {"index": 217795994, "value": "0xc82a3ca4"}, {"index": 208353206, "value": "0xfc72371b"}, {"index": 42483309, "value": "0x4c31b63c"}, {"index": 172547758, "value": "0x0f0d86b4"}, {"index": 148076330, "value": "0x0860e638"}, {"index": 183853158, "value": "0x6016ffc1"}, {"index": 214389424, "value": "0xf1cac2c4"}, {"index": 267488061, "value": "0xb520efed"}, {"index": 169781097, "value": "0x9eac76d6"}, {"index": 184093494, "value": "0xca4ccfe1"}, {"index": 153880993, "value": "0xa07e7108"}, {"index": 84977930, "value": "0xc1c60eac"}, {"index": 46426879, "value": "0xfdd31f2d"}, {"index": 3093825, "value": "0x2225e701"}, {"index": 225364072, "value": "0x71f7fde6"}, {"index": 44593546, "value": "0xb52aed34"}, {"index": 260713159, "value": "0x75a3dc19"}, {"index": 168250303, "value": "0xfe023726"}, {"index": 52384140, "value": "0x308b9e1d"}, {"index": 223401610, "value": "0x6003222c"}, {"index": 45554030, "value": "0xa5cc70a2"}, {"index": 95410555, "value": "0x22c2940f"}, {"index": 175039924, "value": "0xdbcd5db2"}, {"index": 79171087, "value": "0xeed44d67"}, {"index": 267580473, "value": "0xfca48f4c"}, {"index": 24168642, "value": "0x2744926e"}, {"index": 37981670, "value": "0x24f03506"}, {"index": 171551130, "value": "0xa90cad7d"}, {"index": 195559979, "value": "0x47db755e"}, {"index": 204611762, "value": "0x2051b7a9"}, {"index": 140997658, "value": "0x18365d2d"}, {"index": 138925853, "value": "0x69c0d7e2"}, {"index": 86637313, "value": "0x0f1f31c6"}, {"index": 20736778, "value": "0x912e45bf"}, {"index": 219665210, "value": "0x03e4035e"}, {"index": 160430336, "value": "0xa30eb890"}, {"index": 264654675, "value": "0xf65796a0"}, {"index": 8013395, "value": "0x0b97d1e6"}, {"index": 228945585, "value": "0xb0014ea8"}, {"index": 213884386, "value": "0xa02bb1df"}, {"index": 104419827, "value": "0x87476bf2"}, {"index": 44185464, "value": "0xc2f22b36"}, {"index": 142737231, "value": "0x2f0452a5"}, {"index": 99284897, "value": "0xb6d0c4c3"}, {"index": 132475900, "value": "0xc6a765e2"}, {"index": 61861762, "value": "0xf2b8cbd5"}, {"index": 132056166, "value": "0x928e50ad"}, {"index": 262388043, "value": "0x85dbe732"}, {"index": 91878046, "value": "0x2b9ee10c"}, {"index": 117353561, "value": "0x2a1c9124"}, {"index": 124768597, "value": "0x7f9a88e2"}, {"index": 71352993, "value": "0x92da37ea"}, {"index": 190698941, "value": "0x519214cb"}, {"index": 46055428, "value": "0x7869844e"}, {"index": 55281366, "value": "0x19062ee1"}, {"index": 165145231, "value": "0x045f0735"}, {"index": 106810753, "value": "0x09f5e114"}, {"index": 171985651, "value": "0x6082c8d4"}, {"index": 232085256, "value": "0x6bc68181"}, {"index": 159510492, "value": "0x700987d5"}, {"index": 40072060, "value": "0x3bca401d"}, {"index": 209107596, "value": "0x7829f7aa"}, {"index": 39023794, "value": "0xe325295c"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
6943
proto-cuda/packs-ca3-derive/dr736-genesis/kernel.cl
Normal file
6943
proto-cuda/packs-ca3-derive/dr736-genesis/kernel.cl
Normal file
File diff suppressed because it is too large
Load diff
164
proto-cuda/packs-ca3-derive/dr736-genesis/kernel.cu
Normal file
164
proto-cuda/packs-ca3-derive/dr736-genesis/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
7037
proto-cuda/packs-ca3-derive/dr736-genesis/kernel_bound.cl
Normal file
7037
proto-cuda/packs-ca3-derive/dr736-genesis/kernel_bound.cl
Normal file
File diff suppressed because it is too large
Load diff
123
proto-cuda/packs-ca3-derive/dr736-genesis/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca3-derive/dr736-genesis/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
6773
proto-cuda/packs-ca3-derive/dr736-genesis/memhard.h
Normal file
6773
proto-cuda/packs-ca3-derive/dr736-genesis/memhard.h
Normal file
File diff suppressed because it is too large
Load diff
6771
proto-cuda/packs-ca3-derive/dr736-genesis/memhard.metal
Normal file
6771
proto-cuda/packs-ca3-derive/dr736-genesis/memhard.metal
Normal file
File diff suppressed because it is too large
Load diff
72
proto-cuda/packs-ca3-derive/dr736-genesis/program.h
Normal file
72
proto-cuda/packs-ca3-derive/dr736-genesis/program.h
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 2
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x72c1d8048aef9542ull
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "dr736"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 1
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
// Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md, a prototype, NOT class v3): the item
|
||||
// derivation runs the day's drawn program (memhard.h: mh_round_0..8, IGNEUM_DERIVE_LEN instructions each) in place of the mixer.
|
||||
#define IGNEUM_CLASS_DERIVE_LEN 736
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
|
||||
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_DERIVE_LEN 736 // instructions per round program of the day's item-derivation program (Counter ASIC 3.0 item 2; memhard.h mh_round_0..8)
|
||||
#define IGNEUM_DERIVE_ATTEMPT 0
|
||||
#define IGNEUM_DERIVE_FINGERPRINT 0x463535d01511350dull
|
||||
#define IGNEUM_DERIVE_INSTRS_PER_ITEM 6624
|
||||
#define IGNEUM_DERIVE_GPU_OPS_PER_ITEM 10659
|
||||
#define IGNEUM_DERIVE_CHIP_OPS_PER_ITEM 9992
|
||||
#define IGNEUM_DERIVE_MULS_PER_ITEM 1461
|
||||
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
151
proto-cuda/packs-ca3-derive/dr736-genesis/program.json
Normal file
151
proto-cuda/packs-ca3-derive/dr736-genesis/program.json
Normal file
File diff suppressed because one or more lines are too long
109
proto-cuda/packs-ca3-derive/dr736-genesis/program.metal
Normal file
109
proto-cuda/packs-ca3-derive/dr736-genesis/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0
|
||||
r2 = r1 * r1 + r2; // 1
|
||||
r2 = r3 * r2 + r2; // 2
|
||||
r3 = r3 ^ r5; // 3
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 4
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 5
|
||||
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
|
||||
r1 = mulhi(r1, r5); // 8
|
||||
r6 = rotr_var(r6, r3); // 9
|
||||
r3 = r3 | r4; // 10
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 11
|
||||
r0 = mulhi(r0, r4); // 12
|
||||
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
|
||||
r0 = r0 ^ dataset[r4 & MASK]; // 14
|
||||
r2 = r2 - r4; // 15
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 17
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
|
||||
r5 = r5 * r0; // 19
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
|
||||
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
|
||||
r6 = mulhi(r6, r2); // 22
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 23
|
||||
r5 = r5 * r0; // 24
|
||||
r5 = rotl_imm(r5, 19u); // 25
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
|
||||
r0 = r0 ^ r5; // 27
|
||||
r0 = r0 ^ r4; // 28
|
||||
r3 = r3 - r0; // 29
|
||||
r5 = r5 * r1; // 30
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 31
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 32
|
||||
r5 = r5 ^ r6; // 33
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 34
|
||||
r0 = mulhi(r0, r5); // 35
|
||||
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 37
|
||||
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
|
||||
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
|
||||
r2 = r2 ^ r5; // 40
|
||||
r3 = r6 * r3 + r3; // 41
|
||||
r6 = r6 - r7; // 42
|
||||
r7 = r7 ^ r0; // 43
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 44
|
||||
r2 = r2 * r3; // 45
|
||||
r1 = mulhi(r1, r5); // 46
|
||||
r4 = r4 - r3; // 47
|
||||
r2 = rotr_var(r2, r6); // 48
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 49
|
||||
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
|
||||
r0 = r0 * r2; // 51
|
||||
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
|
||||
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
|
||||
r7 = rotl_imm(r7, 14u); // 54
|
||||
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
|
||||
r6 = r6 ^ dataset[r7 & MASK]; // 56
|
||||
r1 = rotr_var(r1, r5); // 57
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 58
|
||||
r6 = r6 ^ dataset[r2 & MASK]; // 59
|
||||
r3 = r5 * r0 + r3; // 60
|
||||
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
|
||||
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
|
||||
r5 = rotl_imm(r5, 19u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca3-derive/dr736-genesis/program_bound.metal
Normal file
111
proto-cuda/packs-ca3-derive/dr736-genesis/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0
|
||||
r2 = r1 * r1 + r2; // 1
|
||||
r2 = r3 * r2 + r2; // 2
|
||||
r3 = r3 ^ r5; // 3
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 4
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 5
|
||||
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
|
||||
r1 = mulhi(r1, r5); // 8
|
||||
r6 = rotr_var(r6, r3); // 9
|
||||
r3 = r3 | r4; // 10
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 11
|
||||
r0 = mulhi(r0, r4); // 12
|
||||
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
|
||||
r0 = r0 ^ dataset[r4 & MASK]; // 14
|
||||
r2 = r2 - r4; // 15
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 17
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
|
||||
r5 = r5 * r0; // 19
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
|
||||
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
|
||||
r6 = mulhi(r6, r2); // 22
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 23
|
||||
r5 = r5 * r0; // 24
|
||||
r5 = rotl_imm(r5, 19u); // 25
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
|
||||
r0 = r0 ^ r5; // 27
|
||||
r0 = r0 ^ r4; // 28
|
||||
r3 = r3 - r0; // 29
|
||||
r5 = r5 * r1; // 30
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 31
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 32
|
||||
r5 = r5 ^ r6; // 33
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 34
|
||||
r0 = mulhi(r0, r5); // 35
|
||||
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 37
|
||||
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
|
||||
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
|
||||
r2 = r2 ^ r5; // 40
|
||||
r3 = r6 * r3 + r3; // 41
|
||||
r6 = r6 - r7; // 42
|
||||
r7 = r7 ^ r0; // 43
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 44
|
||||
r2 = r2 * r3; // 45
|
||||
r1 = mulhi(r1, r5); // 46
|
||||
r4 = r4 - r3; // 47
|
||||
r2 = rotr_var(r2, r6); // 48
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 49
|
||||
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
|
||||
r0 = r0 * r2; // 51
|
||||
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
|
||||
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
|
||||
r7 = rotl_imm(r7, 14u); // 54
|
||||
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
|
||||
r6 = r6 ^ dataset[r7 & MASK]; // 56
|
||||
r1 = rotr_var(r1, r5); // 57
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 58
|
||||
r6 = r6 ^ dataset[r2 & MASK]; // 59
|
||||
r3 = r5 * r0 + r3; // 60
|
||||
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
|
||||
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
|
||||
r5 = rotl_imm(r5, 19u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
57
proto-cuda/packs-ca3-derive/dr736-genesis/vectors.h
Normal file
57
proto-cuda/packs-ca3-derive/dr736-genesis/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0xe23d389f3eea0c83ull, 0xdb20ad461b5a2b40ull, 0x19f7ba3b1f71682cull, 0x85f636966fc9aca0ull, 0x6487ed96bdc27c42ull, 0xc490c3313b8029a3ull, 0xe361eba9844902d4ull, 0xf08e800494c84491ull,
|
||||
0xa8bad93be02c1cacull, 0xe4e1f2c75c891797ull, 0x63b050548aa5d9b3ull, 0x7b8851cd9a064497ull, 0x9873c6fdd34355abull, 0xbe51e127ed02c404ull, 0x1aa1ee5f4c909067ull, 0x422684f3c86c5058ull,
|
||||
0x42c3e7bc0f3578a3ull, 0x3e34aef37a263942ull, 0xcf0a9f5f13bd905cull, 0x44ba872dba90ac37ull, 0xbd590fe07552ad9bull, 0x62b32d43d066161eull, 0x213423a10a13d5d1ull, 0x6bca949b59253f84ull,
|
||||
0xfdb3dba0dcb2951bull, 0x1f871d6c8a19d696ull, 0x413da5a7e4e5f500ull, 0xa448622cdd32bb6dull, 0x5193eba1f8804e2aull, 0x567869c0cfdb19acull, 0x2b2b619c048918dcull, 0x6605db059b381bd9ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0xfdb4b214da8ce292ull, 0xa04a87297fc6a9c0ull, 0x968cecee7a3bcd99ull, 0xa41939b15019ad50ull, 0xffc694d1b04e4da2ull, 0xc262e3973d436fdaull, 0x665a2ce13954b603ull, 0x6ece51e9d9d5d921ull,
|
||||
0x5ff85c609da62800ull, 0xde9a8e850dd6c1e9ull, 0x39e876e19ae70187ull, 0xdb55c10b53c52bf2ull, 0x610845800cd92811ull, 0x696316e0c89c9d52ull, 0x3f18086d783e7d7aull, 0xb0d29c8dfb421fd5ull,
|
||||
0x580543d5977c1138ull, 0xaba172c0bc028b29ull, 0x43ffd7726859240dull, 0xf04dcb8e1b09dcafull, 0xa1ebc5eb5873c4fbull, 0x4141b328723483e0ull, 0xba9fad0940aa4905ull, 0xed0671189b176bd0ull,
|
||||
0x7ed7083b52f320fdull, 0x30c86e3033463bf7ull, 0x8169974907e0f4a3ull, 0x89fd7aad0e41bf03ull, 0x1aab38280fa3ee63ull, 0x6f525aeaa4c8a186ull, 0x5bd2a274c3905d8bull, 0xdb698d03437d74f7ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x534671b1bf5cea36ull, 0xe3f81c66bf210aa3ull, 0xb79ffc3d3cca6731ull, 0x5df44d8e9e3ac008ull, 0x445792668a534b0dull, 0x293f8c07ed96f576ull, 0x5b2e23716649e117ull, 0x55aa3dcc43378f3full,
|
||||
0xfec021749867a3a8ull, 0x3f06d1c7cabc290full, 0x875d22f8b0811415ull, 0x0cd09e82747f0eefull, 0x2bcf3b7cd81a64c8ull, 0x23c180008f4a8677ull, 0x8bae18c724c5810aull, 0xb6a3fe36acf8b5e3ull,
|
||||
0xc145ffcaa45352baull, 0xf1a0f8edaaf82d2dull, 0x205ba1313b3852b1ull, 0x6922192747873b0bull, 0x9937304a65a44a7dull, 0xca589d2fe74a7e01ull, 0xbecb0ddd75ba8dddull, 0x752e223db126522aull,
|
||||
0x92f919e76899e682ull, 0x6c9bf07fddaae5efull, 0xfa7a39ed604236c0ull, 0x1478fb9d0050bc69ull, 0x9e078959178a2797ull, 0x7e6311fd8ff88b13ull, 0x911c33694145c818ull, 0x733b123353f13d21ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0xe73af5cau, 0x45d77effu, 0xb640f499u, 0x351c2ca1u, 0xfed2d16fu, 0x1199bd1du, 0x1db2bbecu, 0x4b1769deu,
|
||||
0x45b535eau, 0x414c389cu, 0xb8cada9cu, 0x32a3eb0fu, 0x405305cau, 0xddbd43e4u, 0xbdd280acu, 0x70358961u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0x7cdbf6b5u;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x0c799a6du, 0xa21ecab8u, 0xcc71d9a3u, 0xb0d7af9cu, 0x1925c929u, 0x457795ffu, 0x4718201au, 0x985cde9eu, 0x4bb9b17eu, 0x677407ddu, 0x6093de87u, 0x2a754135u, 0xbc58f70du, 0xd805eb4bu, 0xf453a94du, 0x5a28b420u, 0xe92fb02bu, 0x8a48d35au, 0x2024c448u, 0x48ac5a95u, 0x8fa7880cu, 0xfff38e1bu, 0x4e98033au, 0xb6a33ee4u, 0xdfb0f94fu, 0x80d99b65u, 0x7d1c8bdeu, 0x06f4c2c4u, 0x05e3527cu, 0x8690f1beu, 0xa392261au, 0xac2293ffu, 0x0d067d65u, 0x7133a61eu, 0x1b115d00u, 0x566e466fu, 0x40e4c461u, 0x39b8ff78u, 0x580c8cafu, 0xbdc4761bu, 0x18e3d94cu, 0x7d742821u, 0xaeb2e213u, 0xbf49ac05u, 0xa607dc56u, 0x225cafb1u, 0xb4f81cf3u, 0x68afbfa5u, 0xe6d35b11u, 0x86050c01u, 0x799c3487u, 0xa90362eau, 0x051a2bbeu, 0xced2c729u, 0xe9429d0bu, 0x5bcdc16cu, 0x0af50eb5u, 0x1493ced7u, 0xcd80ab7eu, 0x4b67919au, 0x633f820au, 0x5fd0275eu, 0x6d20c837u, 0x200a638eu
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
|
||||
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
|
||||
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;
|
||||
36
proto-cuda/packs-ca3-derive/dr736-genesis/vectors.json
Normal file
36
proto-cuda/packs-ca3-derive/dr736-genesis/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-genesis",
|
||||
"day": "2026-10-03",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0xe23d389f3eea0c83", "0xdb20ad461b5a2b40", "0x19f7ba3b1f71682c", "0x85f636966fc9aca0", "0x6487ed96bdc27c42", "0xc490c3313b8029a3", "0xe361eba9844902d4", "0xf08e800494c84491",
|
||||
"0xa8bad93be02c1cac", "0xe4e1f2c75c891797", "0x63b050548aa5d9b3", "0x7b8851cd9a064497", "0x9873c6fdd34355ab", "0xbe51e127ed02c404", "0x1aa1ee5f4c909067", "0x422684f3c86c5058",
|
||||
"0x42c3e7bc0f3578a3", "0x3e34aef37a263942", "0xcf0a9f5f13bd905c", "0x44ba872dba90ac37", "0xbd590fe07552ad9b", "0x62b32d43d066161e", "0x213423a10a13d5d1", "0x6bca949b59253f84",
|
||||
"0xfdb3dba0dcb2951b", "0x1f871d6c8a19d696", "0x413da5a7e4e5f500", "0xa448622cdd32bb6d", "0x5193eba1f8804e2a", "0x567869c0cfdb19ac", "0x2b2b619c048918dc", "0x6605db059b381bd9"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0xfdb4b214da8ce292", "0xa04a87297fc6a9c0", "0x968cecee7a3bcd99", "0xa41939b15019ad50", "0xffc694d1b04e4da2", "0xc262e3973d436fda", "0x665a2ce13954b603", "0x6ece51e9d9d5d921",
|
||||
"0x5ff85c609da62800", "0xde9a8e850dd6c1e9", "0x39e876e19ae70187", "0xdb55c10b53c52bf2", "0x610845800cd92811", "0x696316e0c89c9d52", "0x3f18086d783e7d7a", "0xb0d29c8dfb421fd5",
|
||||
"0x580543d5977c1138", "0xaba172c0bc028b29", "0x43ffd7726859240d", "0xf04dcb8e1b09dcaf", "0xa1ebc5eb5873c4fb", "0x4141b328723483e0", "0xba9fad0940aa4905", "0xed0671189b176bd0",
|
||||
"0x7ed7083b52f320fd", "0x30c86e3033463bf7", "0x8169974907e0f4a3", "0x89fd7aad0e41bf03", "0x1aab38280fa3ee63", "0x6f525aeaa4c8a186", "0x5bd2a274c3905d8b", "0xdb698d03437d74f7"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x534671b1bf5cea36", "0xe3f81c66bf210aa3", "0xb79ffc3d3cca6731", "0x5df44d8e9e3ac008", "0x445792668a534b0d", "0x293f8c07ed96f576", "0x5b2e23716649e117", "0x55aa3dcc43378f3f",
|
||||
"0xfec021749867a3a8", "0x3f06d1c7cabc290f", "0x875d22f8b0811415", "0x0cd09e82747f0eef", "0x2bcf3b7cd81a64c8", "0x23c180008f4a8677", "0x8bae18c724c5810a", "0xb6a3fe36acf8b5e3",
|
||||
"0xc145ffcaa45352ba", "0xf1a0f8edaaf82d2d", "0x205ba1313b3852b1", "0x6922192747873b0b", "0x9937304a65a44a7d", "0xca589d2fe74a7e01", "0xbecb0ddd75ba8ddd", "0x752e223db126522a",
|
||||
"0x92f919e76899e682", "0x6c9bf07fddaae5ef", "0xfa7a39ed604236c0", "0x1478fb9d0050bc69", "0x9e078959178a2797", "0x7e6311fd8ff88b13", "0x911c33694145c818", "0x733b123353f13d21"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0xe73af5ca", "0x45d77eff", "0xb640f499", "0x351c2ca1", "0xfed2d16f", "0x1199bd1d", "0x1db2bbec", "0x4b1769de", "0x45b535ea", "0x414c389c", "0xb8cada9c", "0x32a3eb0f", "0x405305ca", "0xddbd43e4", "0xbdd280ac", "0x70358961"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0x7cdbf6b5",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0x0c799a6d"}, {"index": 217795994, "value": "0xa21ecab8"}, {"index": 208353206, "value": "0xcc71d9a3"}, {"index": 42483309, "value": "0xb0d7af9c"}, {"index": 172547758, "value": "0x1925c929"}, {"index": 148076330, "value": "0x457795ff"}, {"index": 183853158, "value": "0x4718201a"}, {"index": 214389424, "value": "0x985cde9e"}, {"index": 267488061, "value": "0x4bb9b17e"}, {"index": 169781097, "value": "0x677407dd"}, {"index": 184093494, "value": "0x6093de87"}, {"index": 153880993, "value": "0x2a754135"}, {"index": 84977930, "value": "0xbc58f70d"}, {"index": 46426879, "value": "0xd805eb4b"}, {"index": 3093825, "value": "0xf453a94d"}, {"index": 225364072, "value": "0x5a28b420"}, {"index": 44593546, "value": "0xe92fb02b"}, {"index": 260713159, "value": "0x8a48d35a"}, {"index": 168250303, "value": "0x2024c448"}, {"index": 52384140, "value": "0x48ac5a95"}, {"index": 223401610, "value": "0x8fa7880c"}, {"index": 45554030, "value": "0xfff38e1b"}, {"index": 95410555, "value": "0x4e98033a"}, {"index": 175039924, "value": "0xb6a33ee4"}, {"index": 79171087, "value": "0xdfb0f94f"}, {"index": 267580473, "value": "0x80d99b65"}, {"index": 24168642, "value": "0x7d1c8bde"}, {"index": 37981670, "value": "0x06f4c2c4"}, {"index": 171551130, "value": "0x05e3527c"}, {"index": 195559979, "value": "0x8690f1be"}, {"index": 204611762, "value": "0xa392261a"}, {"index": 140997658, "value": "0xac2293ff"}, {"index": 138925853, "value": "0x0d067d65"}, {"index": 86637313, "value": "0x7133a61e"}, {"index": 20736778, "value": "0x1b115d00"}, {"index": 219665210, "value": "0x566e466f"}, {"index": 160430336, "value": "0x40e4c461"}, {"index": 264654675, "value": "0x39b8ff78"}, {"index": 8013395, "value": "0x580c8caf"}, {"index": 228945585, "value": "0xbdc4761b"}, {"index": 213884386, "value": "0x18e3d94c"}, {"index": 104419827, "value": "0x7d742821"}, {"index": 44185464, "value": "0xaeb2e213"}, {"index": 142737231, "value": "0xbf49ac05"}, {"index": 99284897, "value": "0xa607dc56"}, {"index": 132475900, "value": "0x225cafb1"}, {"index": 61861762, "value": "0xb4f81cf3"}, {"index": 132056166, "value": "0x68afbfa5"}, {"index": 262388043, "value": "0xe6d35b11"}, {"index": 91878046, "value": "0x86050c01"}, {"index": 117353561, "value": "0x799c3487"}, {"index": 124768597, "value": "0xa90362ea"}, {"index": 71352993, "value": "0x051a2bbe"}, {"index": 190698941, "value": "0xced2c729"}, {"index": 46055428, "value": "0xe9429d0b"}, {"index": 55281366, "value": "0x5bcdc16c"}, {"index": 165145231, "value": "0x0af50eb5"}, {"index": 106810753, "value": "0x1493ced7"}, {"index": 171985651, "value": "0xcd80ab7e"}, {"index": 232085256, "value": "0x4b67919a"}, {"index": 159510492, "value": "0x633f820a"}, {"index": 40072060, "value": "0x5fd0275e"}, {"index": 209107596, "value": "0x6d20c837"}, {"index": 39023794, "value": "0x200a638e"}],
|
||||
"cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"],
|
||||
"cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"],
|
||||
"cache_fnv1a64": "0x48c4f5bf24166b2e"
|
||||
}
|
||||
118
relay/playbooks/ca3-derive-pc2.ps1
Normal file
118
relay/playbooks/ca3-derive-pc2.ps1
Normal file
|
|
@ -0,0 +1,118 @@
|
|||
# Igneum run job: Counter ASIC 3.0 item 2, the per-day item-derivation program (docs/plans/counter-asic-3-derivation.md),
|
||||
# on PC 2's RTX 5090 (machine 1ccfe586), 6 October 2026. ONE job carries everything (the PC 2 rule of the ca3 brief):
|
||||
# it downloads the packs zip itself from the downloads host (sha256 checked), finds the installed app's
|
||||
# igneum-worker-cuda.exe (NVRTC compiles each pack's own kernel text, so no new worker build is needed), switches the
|
||||
# NVIDIA card off in the app ONLY while the packs run (the card key from the app's settings.json, never /api/state;
|
||||
# restored after with the settings it had), and never quits, restarts or updates the installed app. The numbers this
|
||||
# job is for: per pack, the worker's own `nvrtc .. cache .. dataset .. ms` line (the daily 1 GiB build on the 5090),
|
||||
# the self-test (bit-exactness: cache FNV, dataset head, word [MASK], 64 samples, 96 vector lanes against the Rust CPU
|
||||
# interpreter), the 2^24 fingerprint at base nonce 0 (against the Mac: dr736-genesis 50e3eaa779da4f1e,
|
||||
# dr736-devnet-epoch0 9553f6d5c667205a, mx8-genesis 7c28cfb06c5c65a9, v2-genesis-mh 25f96e7dce90bd4e) and the hash rate.
|
||||
# Packs: dr736-genesis, dr736-devnet-epoch0 (the derivation class), mx8-genesis (the x8 control), v2-genesis-mh (the v2
|
||||
# control). Every result line starts with RESULT. The placeholder __DL_BASE__ is substituted at publish time; the
|
||||
# downloads token never enters the repository.
|
||||
$ErrorActionPreference = 'Continue'
|
||||
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
|
||||
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
|
||||
# everything lives in this job's own folder (the wiped-jobs-folder rule of tools/ci/kit-path-check.sh: nothing of
|
||||
# another job's is reached)
|
||||
$work = Join-Path $env:IGNEUM_JOB_DIR 'ca3-derive'
|
||||
New-Item -ItemType Directory -Force -Path $work | Out-Null
|
||||
$zipUrl = '__DL_BASE__/igneum-ca3-derive-packs.zip'
|
||||
$zipSha = 'aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676'
|
||||
$zip = Join-Path $work 'igneum-ca3-derive-packs.zip'
|
||||
Write-Output "RESULT job ca3-derive-pc2 start $(Stamp) machine $env:COMPUTERNAME"
|
||||
|
||||
# ---- the packs: downloaded by this job, sha256 checked, expanded fresh ----
|
||||
try {
|
||||
[Net.ServicePointManager]::SecurityProtocol = [Net.SecurityProtocolType]::Tls12
|
||||
Invoke-WebRequest -Uri $zipUrl -OutFile $zip -TimeoutSec 120 -Headers @{ 'Cache-Control' = 'no-cache' }
|
||||
} catch { Write-Output ("RESULT error download: " + $_.Exception.Message); exit 2 }
|
||||
$got = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower()
|
||||
if ($got -ne $zipSha) { Write-Output "RESULT error zip sha256 $got expected $zipSha"; exit 2 }
|
||||
Write-Output "RESULT zip sha256 $got size $((Get-Item $zip).Length) OK"
|
||||
$packs = Join-Path $work 'packs-ca3-derive'
|
||||
if (Test-Path $packs) { Remove-Item -Recurse -Force $packs }
|
||||
Expand-Archive -Path $zip -DestinationPath $work -Force
|
||||
$packList = @('v2-genesis-mh', 'mx8-genesis', 'dr736-genesis', 'dr736-devnet-epoch0')
|
||||
foreach ($pk in $packList) {
|
||||
$d = Join-Path $packs $pk
|
||||
if (-not (Test-Path (Join-Path $d 'memhard.h'))) { Write-Output "RESULT error pack $pk missing after extract"; exit 2 }
|
||||
Write-Output ("RESULT pack $pk memhard.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'memhard.h')).Hash.ToLower() + " program.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'program.h')).Hash.ToLower())
|
||||
}
|
||||
|
||||
# ---- the worker: the installed app's igneum-worker-cuda.exe, run in place (its NVRTC DLLs sit beside it) ----
|
||||
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
|
||||
if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe'; exit 2 }
|
||||
$cuda = Join-Path $inst 'igneum-worker-cuda.exe'
|
||||
Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) nvrtc_dlls $((Get-ChildItem $inst -Filter 'nvrtc*.dll').Count)"
|
||||
$help = (& $cuda --help 2>&1 | Out-String)
|
||||
$hasBench = $help -match '--bench'
|
||||
Write-Output "RESULT worker-cuda has --bench: $hasBench"
|
||||
if (-not $hasBench) { Write-Output 'RESULT note the installed worker has no --bench; --check gives the build time and the self-test, no fingerprint and no hash rate' }
|
||||
|
||||
# ---- the app: the NVIDIA card's key and settings from settings.json; switched off through POST api/cards only ----
|
||||
$appDir = $env:IGNEUM_APP_DIR
|
||||
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
|
||||
$urlFile = Join-Path $appDir 'app.url'
|
||||
$url = $null
|
||||
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
|
||||
$sj = Join-Path $appDir 'settings.json'
|
||||
$cardKey = $null; $cardPref = $null
|
||||
if (Test-Path $sj) {
|
||||
try {
|
||||
$settings = Get-Content -LiteralPath $sj -Raw | ConvertFrom-Json
|
||||
if ($settings.cards) {
|
||||
foreach ($p in $settings.cards.PSObject.Properties) { if ($p.Name -like 'nvidia:*') { $cardKey = $p.Name; $cardPref = $p.Value; break } }
|
||||
}
|
||||
} catch { Say ("settings.json: " + $_.Exception.Message) }
|
||||
}
|
||||
if ($cardKey) {
|
||||
Write-Output ("RESULT card " + $cardKey + " enabled=" + $cardPref.enabled + " identities=" + $cardPref.identities + " power_pct=" + $cardPref.power_pct + " (settings.json)")
|
||||
} else {
|
||||
Write-Output 'RESULT card none in settings.json (no nvidia:* entry); the app keeps mining on the card and the numbers carry that load'
|
||||
}
|
||||
$cardOff = $false
|
||||
if ($cardKey -and $url) {
|
||||
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
|
||||
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
|
||||
$body = @{ cards = @(@{ key = $cardKey; enabled = $false; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; $cardOff = $true; Say 'card off requested' } catch { Write-Output ("RESULT error api/cards off: " + $_.Exception.Message) }
|
||||
# wait for the worker process to go, up to 90 s (the process list, not /api/state)
|
||||
$t = 0
|
||||
while ($t -lt 90) {
|
||||
Start-Sleep -Seconds 5; $t += 5
|
||||
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
|
||||
if (-not $w) { break }
|
||||
}
|
||||
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
|
||||
Write-Output ("RESULT card-off " + $cardKey + " after " + $t + " s, worker processes left " + (($w | Measure-Object).Count))
|
||||
Start-Sleep -Seconds 5
|
||||
}
|
||||
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
|
||||
|
||||
# ---- the runs: --check for the build line and the self-test, --bench for the fingerprint and the rate ----
|
||||
foreach ($pk in $packList) {
|
||||
$d = Join-Path $packs $pk
|
||||
Write-Output "RESULT run $pk check start $(Stamp)"
|
||||
& $cuda --check --pack $d 2>&1 | ForEach-Object { "RESULT check $pk $_" }
|
||||
Write-Output "RESULT run $pk check exit $LASTEXITCODE"
|
||||
if ($hasBench) {
|
||||
Write-Output "RESULT run $pk bench start $(Stamp)"
|
||||
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT bench $pk $_" }
|
||||
Write-Output "RESULT run $pk bench exit $LASTEXITCODE"
|
||||
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT bench8 $pk $_" }
|
||||
}
|
||||
}
|
||||
& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
|
||||
|
||||
# ---- restore the card with the settings it had ----
|
||||
if ($cardOff) {
|
||||
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
|
||||
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
|
||||
$en = $true; if ($null -ne $cardPref.enabled) { $en = [bool]$cardPref.enabled }
|
||||
$body = @{ cards = @(@{ key = $cardKey; enabled = $en; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $cardKey + " enabled=" + $en) } catch { Write-Output ("RESULT error card restore " + $cardKey + ": " + $_.Exception.Message) }
|
||||
}
|
||||
Write-Output "RESULT job ca3-derive-pc2 end $(Stamp)"
|
||||
exit 0
|
||||
Loading…
Reference in a new issue