mixer x4 and the cache growth rule (Counter ASIC 2.0, class v3 construction): LoadClass mixer_mult and growth, LoadClass::MX4 (v2 loads, no width roll), memhard::Shape in MixParams, m mixer applications per round with keys round_key(r m + j), Cache::fill_log2, the option C schedule (growth_doublings, cache_log2_words, dataset_log2_words, days_since_genesis) with its test table, day-sized Epoch entries, the three emitters (m loop only for m > 1, v2 text unchanged), program.h and program.json fields, packfile.h mixerMult, packbench and OpenCL host prints, --class mx4 and --days on the CLI; docs/plans/mixer-x4.md design and spec text, docs/analysis/chip-model-v3.md, the 5090 and 9070 XT dataset-build playbooks (measurements and vectors to follow)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
28635b165f
commit
e08909f138
13 changed files with 925 additions and 101 deletions
86
docs/analysis/chip-model-v3.md
Normal file
86
docs/analysis/chip-model-v3.md
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
# The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined
|
||||
|
||||
5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's
|
||||
(`docs/analysis/m16-recompute-attacker-2026-10-05.md`): the strongest chip the plan has priced holds the whole
|
||||
cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations,
|
||||
and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is `50 T / (ops per hash)`. Nothing here
|
||||
is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.
|
||||
|
||||
## 1. Inputs
|
||||
|
||||
| Input | Value | Source |
|
||||
|---|---|---|
|
||||
| Items per hash | 128 (one item per load, 128 loads per hash, median 128.00 distinct) | spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census |
|
||||
| Integer operations per mixer application | about 130 | spec 01 section 1.8.4 |
|
||||
| Mixer applications per item | 9 under v2; 36 under v3 (`m = 4`, `docs/plans/mixer-x4.md`) | `memhard::Shape::mixers_per_item` |
|
||||
| Integer operations per item | 1,170 (v2); 4,680 (v3) | 9 x 130; 36 x 130 |
|
||||
| Integer operations per hash | 149,760 (v2, "150,000"); 599,040 (v3, "600,000") | 128 x the above |
|
||||
| Chip integer budget | 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) | M16 section 3 |
|
||||
| Fixed-function factor | 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) | M16 section 3 |
|
||||
| RTX 5090, version 2 programs, measured | 136.1 MH/s (readwidth, tonight, `docs/plans/read-width.md`, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, `docs/bench-log.md`) | this analysis uses tonight's 136.1 as the denominator and quotes both |
|
||||
| RTX 5090 at w16 (16-byte loads), measured | 139.8 MH/s | readwidth table, tonight (the width stays 4 B: w16 closes nothing) |
|
||||
| Cache mirror, 256 MiB, N5 headline density | 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) | `docs/analysis/sram-mirror.md` revision 2, sections 4 and 5 (`ca2-analysis` e6085c6) |
|
||||
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
|
||||
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
|
||||
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
|
||||
|
||||
## 2. The rows
|
||||
|
||||
Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal
|
||||
silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention
|
||||
("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.
|
||||
|
||||
| Row | Mixer | Ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | With the 3x factor | Equal silicon, SRAM deducted, with the factor |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
|
||||
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
|
||||
| v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
|
||||
| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads) | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.61x or below (owed: the 5090's added-form rate; the hot loads cost it something, the chip nothing) | 1.84x or below | 1.49x |
|
||||
| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.61x or below (owed) | 1.84x or below | 1.45x |
|
||||
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.61x | 1.84x | 1.14x |
|
||||
| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.61x | 1.84x | 0.51x |
|
||||
| v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
|
||||
|
||||
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op
|
||||
weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The
|
||||
width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the
|
||||
chip's cost not at all.
|
||||
|
||||
Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6
|
||||
hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53.
|
||||
Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md`
|
||||
scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. Year 4 and 12 rows: the mirror of
|
||||
`sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of
|
||||
those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area,
|
||||
not a forecast).
|
||||
|
||||
## 3. The margin, plainly
|
||||
|
||||
The combined headline row reads 1.84x with the 3x factor at an equal integer budget, 1.5x with the SRAM area
|
||||
deducted. The claim is "under 2x", and the margin is thin:
|
||||
|
||||
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
|
||||
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
|
||||
- the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
|
||||
- the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from
|
||||
the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if
|
||||
not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.
|
||||
|
||||
What keeps it under 2x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share
|
||||
gave 2.4x, `docs/analysis/scratch-soundness.md` section 3.4; the hot table taxes the DRAM-only chip, not this one;
|
||||
the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in
|
||||
order:
|
||||
|
||||
1. Mixer x8: 0.18x bare and 0.55x with the factor in the M16 table (0.31x and 0.92x against 136.1 here; the M16
|
||||
table's denominator is 229 MH/s); the CPU verifier at 3.3 to 9.6 ms per warp scaled from the version 1 range,
|
||||
at the edge of the 10 ms gate; the measured v3 row of `docs/plans/mixer-x4.md` section 6 is what to scale from
|
||||
now, and whether a 2019-class laptop core (unmeasured, O-1.14) passes 10 ms is what decides it.
|
||||
2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU
|
||||
nothing it was not already paying in cache misses and the chip die area only, so it is the second lever for the
|
||||
chip only through area and the first against a DRAM-only chip.
|
||||
|
||||
## 4. What this does not settle
|
||||
|
||||
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point
|
||||
under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had
|
||||
no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.
|
||||
151
docs/plans/mixer-x4.md
Normal file
151
docs/plans/mixer-x4.md
Normal file
|
|
@ -0,0 +1,151 @@
|
|||
# Mixer x4 and the cache growth rule: the class v3 dataset construction
|
||||
|
||||
5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator
|
||||
under the project lead's delegation at 22:00 UTC (`docs/plans/counter-asic-2-status.md`, "22:00 decided"; the project lead confirms for
|
||||
the public testnet genesis). Branch `ca2-mixer`. Worker: ca2-mixer (cryptographer's lane).
|
||||
|
||||
What this changes, in one line: under program class v3 every mixer application of the dataset item derivation
|
||||
becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the
|
||||
256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two
|
||||
pinned packs re-export without a changed byte (section 5).
|
||||
|
||||
## 1. Why this form
|
||||
|
||||
The recompute attacker of `docs/analysis/m16-recompute-attacker-2026-10-05.md` holds the 256 MiB cache on a die
|
||||
and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8
|
||||
dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day
|
||||
in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier `m` is the one
|
||||
parameter that moves the attacker and leaves the honest hash rate untouched.
|
||||
|
||||
Two shapes give the attacker 4x the operations:
|
||||
|
||||
| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build |
|
||||
|---|---|---|---|---|---|
|
||||
| A, chosen: `m = 4` applications per round, 8 rounds | 36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads |
|
||||
| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches |
|
||||
|
||||
Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the
|
||||
distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up
|
||||
to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if
|
||||
shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6
|
||||
has the measurement; it is under that bound, so B stays unimplemented.
|
||||
|
||||
## 2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)
|
||||
|
||||
### 1.8.5 Item derivation and dataset mapping
|
||||
|
||||
Item `t` (16 words) under mixer multiplier `m` (`m = 1` for program class v2, `m = 4` for class v3; a class
|
||||
parameter, `LoadClass::mixer_mult`):
|
||||
|
||||
```
|
||||
s[i] = K[i] for i in 0..7
|
||||
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
|
||||
for r in 0..7:
|
||||
for j in 0..m-1:
|
||||
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
|
||||
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache
|
||||
s[i] = s[i] XOR cache[line a][i] for i in 0..15
|
||||
for j in 0..m-1:
|
||||
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
|
||||
item(t) = s
|
||||
```
|
||||
|
||||
`M(s, rk)` is the mixer of 1.8.4 with round key `rk`; under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and
|
||||
`9 * 0x9E3779B9`, the version 2 text exactly. The round keys of the `9 m` applications are the first `9 m` values
|
||||
of the version 2 key sequence, all distinct (the sequence is `k * 0x9E3779B9` for `k = 1 .. 9 m`, and
|
||||
`0x9E3779B9` is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every
|
||||
`m` (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. `9 m` mixer
|
||||
applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4).
|
||||
`dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an
|
||||
item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every
|
||||
larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on `C` through
|
||||
the line mask, so the item changes on the day the cache doubles.
|
||||
|
||||
Source: `igneum-pow/src/memhard.rs` (`derive_items`, `round_key_mult`, `Shape`), the emitted `mh_item` of
|
||||
`memhard.h`, `memhard.metal` and `kernel.cl` (`emit.rs`, `emit_memhard_core`: the `m` loop is emitted only for
|
||||
`m > 1`, so every version 2 pack keeps its text).
|
||||
|
||||
### 1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")
|
||||
|
||||
Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, `G x (1 + 86,400 d /
|
||||
(4 x 31,536,000)) = G x (1 + d / 1,460)` for the genesis size `G` and the chain day `d` (DAA days since genesis,
|
||||
section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220
|
||||
(year 28). Rule (Designed, decided 5 October 2026 for class v3):
|
||||
|
||||
```
|
||||
doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2
|
||||
dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
|
||||
cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12
|
||||
```
|
||||
|
||||
Power-of-two sizes only (option (b)), so every load keeps the `src AND MASK` form of 1.14 item 2 and the cache
|
||||
line index keeps `s[0] AND mask`. The cache doubles exactly when the dataset doubles ("option C",
|
||||
`docs/analysis/sram-mirror.md` section 7): the cache's job is to stay above any GPU's last-level cache and that
|
||||
needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).
|
||||
|
||||
`d` is `day_index(header.timestamp) - day_index(genesis.timestamp)` with `day_index = timestamp_ms / 86,400,000`
|
||||
(`bind::day_index`, the interim day rule), clamped at 0 (`memhard::days_since_genesis`). Under class v2 nothing
|
||||
grows: the cache is 2^26 words and the dataset the genesis size on every day.
|
||||
|
||||
| Chain day `d` | Years | `doublings` | Cache words | Cache | Dataset words (devnet `D_0 = 28`) | Dataset (designed `D_0 = 29`) | Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s |
|
||||
| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s |
|
||||
| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s |
|
||||
| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s |
|
||||
| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s |
|
||||
|
||||
Test: `memhard::tests::growth_schedule_table` pins every row and the day before each step. The devnet pack
|
||||
`igneum-devnet-v4-epoch0` is day 20,730 of the Unix count against genesis day 20,729, `d = 1`, so every existing
|
||||
size and vector stands.
|
||||
|
||||
Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4
|
||||
and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card
|
||||
holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until
|
||||
year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built,
|
||||
the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps;
|
||||
option (a) would have faded a 4 GiB card out in year 4 instead of year 4.
|
||||
|
||||
## 3. Interfaces
|
||||
|
||||
| Item | Where | Note |
|
||||
|---|---|---|
|
||||
| `LoadClass { mixer_mult: u8, growth: bool }`, `LoadClass::MX4` ("mx4"), `with_mixer(m, growth)`, `v2_loads()`, `takes_width_roll()` | `generator.rs` | a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs |
|
||||
| `Shape { mixer_mult, cache_log2_words }`, `Shape::for_class_day(class, d)`, `MixParams.shape`, `Cache::fill_log2(key, log2)`, `round_key_mult(r, j, m)` | `memhard.rs` | `Shape::V2` is version 2 |
|
||||
| `growth_doublings(d)`, `cache_log2_words(d)`, `dataset_log2_words(D_0, d)`, `days_since_genesis(day, genesis_day)` | `memhard.rs` | the schedule, one function and its two sizes |
|
||||
| `DatasetSource::{new_shape, from_key_shape, shape}`, `Epoch::new_class_day`, `Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0)` | `verify.rs` | the day-sized entries; the v2 entries are unchanged and build the v2 shape |
|
||||
| `IGNEUM_MIXER_MULT`, `IGNEUM_CLASS_MIXER_MULT`, `IGNEUM_CACHE_GROWTH` in program.h; `"mixer_mult"`, `"cache_growth"`, the `"item"` string in program.json | `emit.rs` | written only for a class with `m != 1` or growth, so v2 packs do not change |
|
||||
| `packfile.h` `mixerMult`; `packbench` and the OpenCL host print the multiplier and the cache size | the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from `IGNEUM_CACHE_LOG2_WORDS` already (packbench.swift line 56, host.cu line 65, host.c line 1966) |
|
||||
| `igneum-pow --class mx4 [--days d]` on every command | `main.rs` | `--days` sizes the cache for a growth class |
|
||||
|
||||
Under the ca2-v3 seam (`ProgramClass::V3`, `V3_CLASS`), the integration sets `V3_CLASS = LoadClass::MX4`; the
|
||||
chain's day-sized dataset needs the day index, so `Epoch::chain_dataset(day, class)` builds the genesis-size
|
||||
cache and a `chain_dataset_day(day_bytes, class, d, D_0)` beside it is the growth entry (section 9, owed to the
|
||||
node agent).
|
||||
|
||||
## 4. Vectors (class v3, `proto-cuda/packs-ca2-mixer/`)
|
||||
|
||||
Filled in section 6 from the exported packs: `mx4-genesis` (seed `igneum-genesis`, day `2026-10-03`, 2^28 words,
|
||||
2^26-word cache) and `mx4-devnet-epoch0` (the devnet genesis hash as the epoch seed, day bytes of 2026-10-04).
|
||||
|
||||
## 5. The v2 path is byte-identical
|
||||
|
||||
`cargo test --test packs` regenerates every file of `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` from
|
||||
`program.json` and compares byte for byte (`emitted_sources_match_all_packs`, `export_pack_matches_all_packs`);
|
||||
section 6 also records a fresh `igneum-pow export` of both packs diffed against the checked-in directories.
|
||||
|
||||
## 6. Measurements
|
||||
|
||||
Filled as they land. Every row names the machine, the date, the command and the lock mode.
|
||||
|
||||
## 7. The chip model
|
||||
|
||||
`docs/analysis/chip-model-v3.md`.
|
||||
|
||||
## 8. What is unverified
|
||||
|
||||
Filled at the end.
|
||||
|
||||
## 9. Owed
|
||||
|
||||
Filled at the end.
|
||||
|
|
@ -11,8 +11,8 @@
|
|||
|
||||
use crate::generator::{Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
|
||||
use crate::memhard::{
|
||||
MixParams, CACHE_LINES_PER_SEGMENT, CACHE_LINE_MASK, CACHE_LOG2_WORDS, CACHE_SEGMENTS, CACHE_SEGMENT_LOG2_LINES,
|
||||
CACHE_TAG, CACHE_WORDS, CHACHA_ROUNDS, CHACHA_SIGMA, ITEM_ROUNDS,
|
||||
MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, CHACHA_ROUNDS, CHACHA_SIGMA,
|
||||
ITEM_ROUNDS,
|
||||
};
|
||||
use crate::seed::SplitMix64;
|
||||
use crate::verify::{DatasetMode, DatasetSource, Epoch, FOLD_MUL, FOLD_ROT};
|
||||
|
|
@ -100,12 +100,25 @@ fn class_header_lines(p: &Program) -> String {
|
|||
return String::new();
|
||||
}
|
||||
let mut s = String::new();
|
||||
s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads
|
||||
if p.class.v2_loads() {
|
||||
s.push_str("// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
");
|
||||
s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x.
|
||||
s.push_str("// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
");
|
||||
} else {
|
||||
s.push_str("// Read-width experiment (5 October 2026, docs/plans/read-width.md): NOT the lottery hash. A load of W words reads
|
||||
");
|
||||
s.push_str("// the W-word-aligned address and folds every word into dst: x = dst ^ w[0]; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x.
|
||||
");
|
||||
}
|
||||
s.push_str(&format!("#define IGNEUM_LOAD_CLASS {}
|
||||
", jstr(&p.class.name())));
|
||||
if p.class.mixer_mult != 1 || p.class.growth {
|
||||
s.push_str(&format!("#define IGNEUM_CLASS_MIXER_MULT {}
|
||||
", p.class.mixer_mult));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_GROWTH {} // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
", p.class.growth as u8));
|
||||
}
|
||||
s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {}
|
||||
", p.class.load_slots));
|
||||
s.push_str(&format!("#define IGNEUM_LOAD_MIX {{ {}, {}, {} }}
|
||||
|
|
@ -257,12 +270,19 @@ pub enum LoadSource<'a> {
|
|||
InlineMemhard(&'a MixParams),
|
||||
}
|
||||
|
||||
fn log2_segments() -> usize {
|
||||
CACHE_SEGMENTS.trailing_zeros() as usize
|
||||
fn log2_segments(shape: &Shape) -> usize {
|
||||
shape.log2_segments() as usize
|
||||
}
|
||||
|
||||
/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal.
|
||||
/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal: the mixer constants of
|
||||
/// the day and, from `mp.shape`, the cache size and the mixer multiplier `m`. Under `m = 1` the text is version
|
||||
/// 2's byte for byte; under `m > 1` the item loop applies `mh_mixer` `m` times per round with the keys
|
||||
/// `round_key(r m + j)` (spec 01 section 1.8.5 under class v3).
|
||||
pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
|
||||
let shape = &mp.shape;
|
||||
let m = shape.mixer_mult;
|
||||
let cache_log2_words = shape.cache_log2_words;
|
||||
let cache_line_mask = shape.cache_line_mask();
|
||||
let (u, fn_, cptr, wptr, lptr, lcptr) = match dialect {
|
||||
CoreDialect::Metal => {
|
||||
("uint", "inline", "device const uint*", "device uint*", "thread uint*", "const thread uint*")
|
||||
|
|
@ -274,18 +294,23 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
|
|||
};
|
||||
let k = &mp.key;
|
||||
let r = &mp.rot;
|
||||
let m = &mp.mul;
|
||||
let mul = &mp.mul;
|
||||
let c = &mp.rc;
|
||||
let mut s = String::with_capacity(6000);
|
||||
s.push_str(&format!(
|
||||
"// Memory-hard dataset core (MEMHARD.md). Cache: 2^{} words in 2^{} segments of {} chained ChaCha{} lines.\n",
|
||||
CACHE_LOG2_WORDS,
|
||||
log2_segments(),
|
||||
cache_log2_words,
|
||||
log2_segments(shape),
|
||||
CACHE_LINES_PER_SEGMENT,
|
||||
CHACHA_ROUNDS
|
||||
));
|
||||
s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n");
|
||||
s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(CACHE_LINE_MASK)));
|
||||
if m == 1 {
|
||||
s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n");
|
||||
} else {
|
||||
s.push_str(&format!("// Item: 8 rounds of {m} x seed-parameterised mixer + one 64-byte cache read, then {m} x final mixer (class v3, mixer multiplier {m},\n"));
|
||||
s.push_str("// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.\n");
|
||||
}
|
||||
s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(cache_line_mask)));
|
||||
s.push_str(&format!("#define MH_SEGMENT_LINES {}u\n", CACHE_LINES_PER_SEGMENT));
|
||||
s.push_str("#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }\n");
|
||||
s.push_str(&format!(
|
||||
|
|
@ -345,7 +370,7 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
|
|||
s.push_str("// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.\n");
|
||||
s.push_str(&format!("{fn_} void mh_mixer({lptr} s, {u} rk) {{\n"));
|
||||
for i in 0..16 {
|
||||
s.push_str(&format!(" s[{i}] = (s[{i}] ^ ({} + rk)) * {};\n", hex(c[i]), hex(m[i])));
|
||||
s.push_str(&format!(" s[{i}] = (s[{i}] ^ ({} + rk)) * {};\n", hex(c[i]), hex(mul[i])));
|
||||
}
|
||||
let col = (0..4).map(|i| format!("{}u", r[i])).collect::<Vec<_>>().join(", ");
|
||||
let dia = (4..8).map(|i| format!("{}u", r[i])).collect::<Vec<_>>().join(", ");
|
||||
|
|
@ -355,22 +380,39 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
|
|||
s.push_str(&format!(" MH_QR(s[2], s[7], s[8], s[13], {dia}) MH_QR(s[3], s[4], s[9], s[14], {dia})\n"));
|
||||
s.push_str("}\n");
|
||||
s.push('\n');
|
||||
s.push_str(&format!(
|
||||
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of mixer + cache line s[0] & mask; final mixer.\n"
|
||||
));
|
||||
if m == 1 {
|
||||
s.push_str(&format!(
|
||||
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of mixer + cache line s[0] & mask; final mixer.\n"
|
||||
));
|
||||
} else {
|
||||
s.push_str(&format!(
|
||||
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of {m} x mixer + cache line s[0] & mask; {m} x final mixer.\n"
|
||||
));
|
||||
}
|
||||
s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n"));
|
||||
for i in 0..8 {
|
||||
s.push_str(&format!(" s[{i}] = {};\n", hex(k[i])));
|
||||
}
|
||||
for i in 0..8 {
|
||||
s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(m[i]), hex(c[i])));
|
||||
s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(mul[i]), hex(c[i])));
|
||||
}
|
||||
s.push_str(&format!(" for ({u} r = 0u; r < {ITEM_ROUNDS}u; ++r) {{\n"));
|
||||
s.push_str(" mh_mixer(s, 0x9E3779B9u * (r + 1u));\n");
|
||||
if m == 1 {
|
||||
s.push_str(" mh_mixer(s, 0x9E3779B9u * (r + 1u));\n");
|
||||
} else {
|
||||
s.push_str(&format!(" for ({u} j = 0u; j < {m}u; ++j) mh_mixer(s, 0x9E3779B9u * (r * {m}u + j + 1u));\n"));
|
||||
}
|
||||
s.push_str(&format!(" {cptr} line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);\n"));
|
||||
s.push_str(&format!(" for ({u} i = 0u; i < 16u; ++i) s[i] ^= line[i];\n"));
|
||||
s.push_str(" }\n");
|
||||
s.push_str(&format!(" mh_mixer(s, 0x9E3779B9u * {}u);\n", ITEM_ROUNDS + 1));
|
||||
if m == 1 {
|
||||
s.push_str(&format!(" mh_mixer(s, 0x9E3779B9u * {}u);\n", ITEM_ROUNDS + 1));
|
||||
} else {
|
||||
s.push_str(&format!(
|
||||
" for ({u} j = 0u; j < {m}u; ++j) mh_mixer(s, 0x9E3779B9u * ({}u + j + 1u));\n",
|
||||
ITEM_ROUNDS as u32 * m
|
||||
));
|
||||
}
|
||||
s.push_str("}\n");
|
||||
s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n");
|
||||
s.push_str(&format!(
|
||||
|
|
@ -386,7 +428,7 @@ pub fn metal_memhard(mp: &MixParams) -> String {
|
|||
s.push_str("using namespace metal;\n");
|
||||
s.push_str(&emit_memhard_core(mp, CoreDialect::Metal));
|
||||
s.push('\n');
|
||||
s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments()));
|
||||
s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments(&mp.shape)));
|
||||
s.push_str(
|
||||
"kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {\n",
|
||||
);
|
||||
|
|
@ -1116,10 +1158,13 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
s.push_str(&format!("#define IGNEUM_SEEDW_INIT {{ {} }}\n", join_hex(&p.seed)));
|
||||
if let Some(mp) = memhard {
|
||||
s.push_str(&format!("#define IGNEUM_KEY_INIT {{ {} }}\n", join_hex(&mp.key)));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_LOG2_WORDS {CACHE_LOG2_WORDS}\n"));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_LOG2_WORDS {}\n", mp.shape.cache_log2_words));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_SEGMENT_LOG2_LINES {CACHE_SEGMENT_LOG2_LINES}\n"));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_SEGMENTS {CACHE_SEGMENTS}u\n"));
|
||||
s.push_str(&format!("#define IGNEUM_CACHE_SEGMENTS {}u\n", mp.shape.cache_segments()));
|
||||
s.push_str(&format!("#define IGNEUM_ITEM_ROUNDS {ITEM_ROUNDS}\n"));
|
||||
if mp.shape.mixer_mult != 1 {
|
||||
s.push_str(&format!("#define IGNEUM_MIXER_MULT {} // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)\n", mp.shape.mixer_mult));
|
||||
}
|
||||
s.push_str(&format!(
|
||||
"#define IGNEUM_MIX_ROT_INIT {{ {} }}\n",
|
||||
mp.rot.iter().map(|r| format!("{r}u")).collect::<Vec<_>>().join(", ")
|
||||
|
|
@ -1188,6 +1233,8 @@ pub struct PackVectors {
|
|||
pub cache_last: Vec<u32>,
|
||||
/// FNV-1a 64 over the whole cache (memory-hard only)
|
||||
pub cache_fnv: u64,
|
||||
/// The cache is 2^cache_log2_words words (memory-hard only; 26 under version 2)
|
||||
pub cache_log2_words: u32,
|
||||
}
|
||||
|
||||
/// The base nonces of the three vector warps every pack carries.
|
||||
|
|
@ -1249,7 +1296,8 @@ pub fn vectors_header(
|
|||
s.push_str("};\n");
|
||||
if memhard {
|
||||
s.push_str(&format!(
|
||||
"// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^{CACHE_LOG2_WORDS} words.\n"
|
||||
"// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^{} words.\n",
|
||||
v.cache_log2_words
|
||||
));
|
||||
s.push_str("static const uint32_t IGNEUM_CACHE_HEAD[16] = {\n");
|
||||
s.push_str(&format!(" {},\n", join_hex(&v.cache_head[..8])));
|
||||
|
|
@ -1300,6 +1348,11 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
if !p.class.is_v2() {
|
||||
let c = p.width_counts();
|
||||
s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name())));
|
||||
if p.class.mixer_mult != 1 || p.class.growth {
|
||||
s.push_str(&format!(" \"mixer_mult\": {},\n", p.class.mixer_mult));
|
||||
s.push_str(&format!(" \"cache_growth\": {},\n", p.class.growth));
|
||||
s.push_str(&format!(" \"mixer\": \"class v3 (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): every mixer application of the item derivation is {} applications with round keys (r * {} + j + 1) * 0x9E3779B9, the 8 dependent cache reads per item unchanged; cache growth rule option C: cache words = 2^(26 + doublings(day)), dataset words = 2^(genesis_log2 + doublings(day)), doublings(day) = floor(log2(1 + day / 1460)) for day = days since genesis\",\n", p.class.mixer_mult, p.class.mixer_mult));
|
||||
}
|
||||
s.push_str(&format!(" \"load_slots\": {},\n", p.class.load_slots));
|
||||
s.push_str(&format!(" \"load_mix_percent_4_16_64\": [{}, {}, {}],\n", p.class.mix[0], p.class.mix[1], p.class.mix[2]));
|
||||
s.push_str(&format!(" \"load_width_counts_4_16_64\": [{}, {}, {}],\n", c[0], c[1], c[2]));
|
||||
|
|
@ -1351,9 +1404,12 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
s.push_str(" \"spec\": \"proto-metal/MEMHARD.md\",\n");
|
||||
s.push_str(&format!(" \"key\": [{}],\n", join_jhex(&mp.key)));
|
||||
s.push_str(" \"key_derivation\": \"the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]\",\n");
|
||||
let shape = &mp.shape;
|
||||
s.push_str(&format!(
|
||||
" \"cache\": {{\"log2_words\": {CACHE_LOG2_WORDS}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {CACHE_SEGMENTS}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n",
|
||||
CACHE_WORDS as u64 * 4,
|
||||
" \"cache\": {{\"log2_words\": {}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n",
|
||||
shape.cache_log2_words,
|
||||
shape.cache_words() as u64 * 4,
|
||||
shape.cache_segments(),
|
||||
join_jhex(&CHACHA_SIGMA),
|
||||
join_jhex(&CACHE_TAG)
|
||||
));
|
||||
|
|
@ -1364,11 +1420,23 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
join_jhex(&mp.rc)
|
||||
));
|
||||
// The Swift writes jhex(cacheLineMask) here, which breaks the JSON. We write the bare literal.
|
||||
s.push_str(&format!(
|
||||
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n",
|
||||
ITEM_ROUNDS - 1,
|
||||
CACHE_LINE_MASK
|
||||
));
|
||||
if shape.mixer_mult == 1 {
|
||||
s.push_str(&format!(
|
||||
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n",
|
||||
ITEM_ROUNDS - 1,
|
||||
shape.cache_line_mask()
|
||||
));
|
||||
} else {
|
||||
let m = shape.mixer_mult;
|
||||
s.push_str(&format!(
|
||||
" \"mixer_mult\": {m},\n \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: for j in 0..{}: s = M(s, rk = (r * {m} + j + 1) * 0x9E3779B9); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then for j in 0..{}: s = M(s, rk = ({} + j + 1) * 0x9E3779B9); item(t) = s\",\n",
|
||||
ITEM_ROUNDS - 1,
|
||||
m - 1,
|
||||
shape.cache_line_mask(),
|
||||
m - 1,
|
||||
ITEM_ROUNDS as u32 * m
|
||||
));
|
||||
}
|
||||
s.push_str(" \"word\": \"dataset[w] = item(w >> 4)[w & 15]\"\n");
|
||||
} else {
|
||||
s.push_str(" \"mode\": \"closed-form\",\n");
|
||||
|
|
@ -1507,6 +1575,7 @@ pub fn export_pack(epoch: &Epoch, day: &str, source: &str) -> Pack {
|
|||
v.cache_head = w[..16].to_vec();
|
||||
v.cache_last = w[w.len() - 16..].to_vec();
|
||||
v.cache_fnv = m.cache.fnv1a64();
|
||||
v.cache_log2_words = m.shape().cache_log2_words;
|
||||
}
|
||||
let is_mh = memhard.is_some();
|
||||
let mut files = vec![
|
||||
|
|
|
|||
|
|
@ -181,6 +181,13 @@ pub struct LoadClass {
|
|||
/// Variant 5: the scratch per warp in KiB (32 or 128; the whole working set of a card at full occupancy must
|
||||
/// stay under 6 GB, coordinator's cap of 5 October 2026). 0 for every other class.
|
||||
pub scratch_kb: u8,
|
||||
/// Mixer cost multiplier `m` of the dataset item derivation (Counter ASIC 2.0, M16, decided 5 October 2026 for
|
||||
/// class v3): every mixer application of spec 01 section 1.8.5 becomes `m` applications with distinct round
|
||||
/// keys, the 8 dependent cache reads per item unchanged (`memhard::derive_items`). 1 for version 2, 4 for v3.
|
||||
pub mixer_mult: u8,
|
||||
/// Cache growth rule, option C (`memhard::growth_doublings`): the cache doubles when the dataset doubles. `false`
|
||||
/// for version 2 (the cache is 2^26 words on every day), `true` for v3.
|
||||
pub growth: bool,
|
||||
}
|
||||
|
||||
/// Scratch geometry (variant 5): 16-byte slots, lane-major, 32 lanes per warp; `scratch_kb` KiB per warp gives
|
||||
|
|
@ -205,20 +212,27 @@ impl LoadClass {
|
|||
|
||||
impl LoadClass {
|
||||
/// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash.
|
||||
pub const V2: LoadClass = LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 };
|
||||
pub const V2: LoadClass =
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false };
|
||||
|
||||
/// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`):
|
||||
/// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the
|
||||
/// mixer applied 4 times per round, and the cache growth rule. Name "mx4".
|
||||
pub const MX4: LoadClass =
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true };
|
||||
|
||||
/// A fixed width (1, 4 or 16 words) with `load_slots` loads per program.
|
||||
pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass {
|
||||
let mut mix = [0u8; 3];
|
||||
let i = WIDTH_WORDS.iter().position(|&w| w == width_words).expect("width must be 1, 4 or 16 words");
|
||||
mix[i] = 100;
|
||||
LoadClass { mix, load_slots, scratch: None, scratch_kb: 0 }
|
||||
LoadClass { mix, load_slots, ..LoadClass::V2 }
|
||||
}
|
||||
|
||||
/// Per-load width drawn from `mix` (percent for 4, 16, 64 bytes), 16 loads per program.
|
||||
pub fn mixed(mix: [u8; 3]) -> LoadClass {
|
||||
assert_eq!(mix.iter().map(|&m| m as u32).sum::<u32>(), 100, "the mix must sum to 100");
|
||||
LoadClass { mix, load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 }
|
||||
LoadClass { mix, ..LoadClass::V2 }
|
||||
}
|
||||
|
||||
/// Variant 5: version 2 widths, 16 memory operations of which `k` are scratch read-modify-writes into a
|
||||
|
|
@ -226,11 +240,61 @@ impl LoadClass {
|
|||
pub fn scratch(k: u8, kb: u8) -> LoadClass {
|
||||
assert!(k as usize <= LOAD_SLOTS);
|
||||
assert!(kb.is_power_of_two() && kb <= 128, "scratch per warp must be a power of two up to 128 KiB");
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: Some(k), scratch_kb: kb }
|
||||
LoadClass { scratch: Some(k), scratch_kb: kb, ..LoadClass::V2 }
|
||||
}
|
||||
|
||||
/// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp).
|
||||
/// This class with the mixer multiplier `m` (1, 2, 4, 8 or 16) and the cache growth rule on or off.
|
||||
pub fn with_mixer(self, mixer_mult: u8, growth: bool) -> LoadClass {
|
||||
assert!(mixer_mult >= 1 && mixer_mult <= 16 && mixer_mult.is_power_of_two(), "mixer multiplier must be 1, 2, 4, 8 or 16");
|
||||
LoadClass { mixer_mult, growth, ..self }
|
||||
}
|
||||
|
||||
/// The mixer multiplier as the item derivation uses it.
|
||||
pub fn mixer_mult(&self) -> u32 {
|
||||
self.mixer_mult as u32
|
||||
}
|
||||
|
||||
/// Whether the loads of this class are version 2's: 16 one-word loads, no scratch. Such a class takes no width
|
||||
/// roll, so its program stream is the version 2 stream draw for draw (the mixer and the cache are properties of
|
||||
/// the dataset, not of the program).
|
||||
pub fn v2_loads(&self) -> bool {
|
||||
self.mix == [100, 0, 0] && self.load_slots as usize == LOAD_SLOTS && self.scratch.is_none()
|
||||
}
|
||||
|
||||
/// Whether every instruction takes the tenth draw (the width roll): every class whose loads are not version 2's.
|
||||
pub fn takes_width_roll(&self) -> bool {
|
||||
!self.v2_loads()
|
||||
}
|
||||
|
||||
/// Parse "p4,p16,p64" or one of the names of [`LoadClass::name`] ("scr4k32": 4 scratch ops, 32 KiB per warp;
|
||||
/// "mx4": the v3 construction; a trailing "m<mult>" and "g" set the mixer multiplier and the growth rule on any
|
||||
/// load class, "w16m4g" for example).
|
||||
pub fn parse(s: &str) -> Option<LoadClass> {
|
||||
if s == "mx4" {
|
||||
return Some(LoadClass::MX4);
|
||||
}
|
||||
// the mixer suffix: "...m<mult>" then an optional "g"
|
||||
let (s, growth) = match s.strip_suffix('g') {
|
||||
Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true),
|
||||
_ => (s, false),
|
||||
};
|
||||
if let Some((base, digits)) = s.rsplit_once('m') {
|
||||
if !digits.is_empty() && digits.bytes().all(|b| b.is_ascii_digit()) && !base.is_empty() && !base.ends_with(',') {
|
||||
let mult: u8 = digits.parse().ok()?;
|
||||
if mult == 0 || mult > 16 || !mult.is_power_of_two() {
|
||||
return None;
|
||||
}
|
||||
return Some(LoadClass::parse_loads(base)?.with_mixer(mult, growth));
|
||||
}
|
||||
}
|
||||
if growth {
|
||||
return None;
|
||||
}
|
||||
LoadClass::parse_loads(s)
|
||||
}
|
||||
|
||||
/// The load part of a class name (no mixer suffix).
|
||||
fn parse_loads(s: &str) -> Option<LoadClass> {
|
||||
if let Some(rest) = s.strip_prefix("scr") {
|
||||
let (k, kb) = rest.split_once('k')?;
|
||||
let k: u8 = k.parse().ok()?;
|
||||
|
|
@ -260,7 +324,7 @@ impl LoadClass {
|
|||
if slots == 0 || slots as usize >= INSTR_COUNT {
|
||||
return None;
|
||||
}
|
||||
Some(LoadClass { mix, load_slots: slots, scratch: None, scratch_kb: 0 })
|
||||
Some(LoadClass { mix, load_slots: slots, ..LoadClass::V2 })
|
||||
}
|
||||
|
||||
/// Scratch read-modify-writes per program (0 without a scratch).
|
||||
|
|
@ -272,25 +336,41 @@ impl LoadClass {
|
|||
*self == LoadClass::V2
|
||||
}
|
||||
|
||||
/// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4".
|
||||
/// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4k32"; "mx4" for the v3 construction;
|
||||
/// any other mixer setting appends "m<mult>" and, with the growth rule, "g" ("v2m2", "w16m4g").
|
||||
pub fn name(&self) -> String {
|
||||
if self.is_v2() {
|
||||
return "v2".to_string();
|
||||
}
|
||||
if let Some(k) = self.scratch {
|
||||
return format!("scr{k}k{}", self.scratch_kb);
|
||||
if *self == LoadClass::MX4 {
|
||||
return "mx4".to_string();
|
||||
}
|
||||
let base = match self.mix {
|
||||
[100, 0, 0] => "w4".to_string(),
|
||||
[0, 100, 0] => "w16".to_string(),
|
||||
[0, 0, 100] => "w64".to_string(),
|
||||
[a, b, c] => format!("mix{a}-{b}-{c}"),
|
||||
};
|
||||
if self.load_slots as usize == LOAD_SLOTS {
|
||||
base
|
||||
let loads = LoadClass { mixer_mult: 1, growth: false, ..*self };
|
||||
let base = if loads.is_v2() {
|
||||
"v2".to_string()
|
||||
} else if let Some(k) = self.scratch {
|
||||
format!("scr{k}k{}", self.scratch_kb)
|
||||
} else {
|
||||
format!("{base}x{}", self.load_slots)
|
||||
let base = match self.mix {
|
||||
[100, 0, 0] => "w4".to_string(),
|
||||
[0, 100, 0] => "w16".to_string(),
|
||||
[0, 0, 100] => "w64".to_string(),
|
||||
[a, b, c] => format!("mix{a}-{b}-{c}"),
|
||||
};
|
||||
if self.load_slots as usize == LOAD_SLOTS {
|
||||
base
|
||||
} else {
|
||||
format!("{base}x{}", self.load_slots)
|
||||
}
|
||||
};
|
||||
let mut s = base;
|
||||
if self.mixer_mult != 1 || self.growth {
|
||||
s.push_str(&format!("m{}", self.mixer_mult));
|
||||
}
|
||||
if self.growth {
|
||||
s.push('g');
|
||||
}
|
||||
s
|
||||
}
|
||||
|
||||
/// The width in words of a load whose width roll (0..99) is `roll`: the first entry of the mix whose cumulative
|
||||
|
|
@ -333,12 +413,11 @@ pub enum ProgramClass {
|
|||
V3,
|
||||
}
|
||||
|
||||
/// PLACEHOLDER (ca2-node worker, 5 October 2026): the load class of program class v3 is set to w16 (16 loads of
|
||||
/// four words, `LoadClass::fixed(4, 16)`) so the node, the miner, the packs and the fast-time gate can be built and
|
||||
/// run before the Counter ASIC 2.0 measurements decide the width, the per-load mix and the scratch share
|
||||
/// (`counter-asic-2-rollout.md` section 6). The integration agent on branch ca2-v3 replaces this constant with the
|
||||
/// decided class; nothing else in the seam names the class, so it is one edit.
|
||||
pub const V3_CLASS: LoadClass = LoadClass { mix: [0, 100, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0 };
|
||||
/// The load class of program class v3, decided 5 October 2026 (Counter ASIC 2.0, `docs/plans/counter-asic-2-status.md`
|
||||
/// "22:00 decided", `docs/plans/mixer-x4.md`): [`LoadClass::MX4`], version 2 loads (the width stays 4 bytes, the
|
||||
/// per-load mix and the scratch share are out), the mixer applied 4 times per round and the cache growth rule. The
|
||||
/// placeholder of the seam (w16) is replaced here; nothing else in the seam names the class.
|
||||
pub const V3_CLASS: LoadClass = LoadClass::MX4;
|
||||
|
||||
impl ProgramClass {
|
||||
/// The load class this program class draws from.
|
||||
|
|
@ -494,6 +573,14 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L
|
|||
b.push(k);
|
||||
b.push(class.scratch_kb);
|
||||
}
|
||||
if class.mixer_mult != 1 || class.growth {
|
||||
// Counter ASIC 2.0: the mixer multiplier and the growth rule are part of the construction, so a program of
|
||||
// the same seed under a different mixer carries a different id (under the v3 seam the id is
|
||||
// program_id(3, seed, attempt) and this branch is not taken)
|
||||
b.extend_from_slice(b"mixer/");
|
||||
b.push(class.mixer_mult);
|
||||
b.push(class.growth as u8);
|
||||
}
|
||||
fnv1a64(&b)
|
||||
}
|
||||
|
||||
|
|
@ -625,7 +712,8 @@ pub fn candidate_from_words_class(
|
|||
let rot = 1 + rng.below(31) as u32;
|
||||
let bit = rng.below(32);
|
||||
let mask = 1u8 << rng.below(5);
|
||||
let width = if class.is_v2() { 1 } else { class.width_for_roll(rng.below(100)) };
|
||||
// Version 2 loads take no width roll, so a mixer class with version 2 loads draws the version 2 program
|
||||
let width = if class.takes_width_roll() { class.width_for_roll(rng.below(100)) } else { 1 };
|
||||
let width = if op == Op::Load { width } else { 1 };
|
||||
if op.is_load() {
|
||||
fresh[src as usize] = false;
|
||||
|
|
|
|||
|
|
@ -32,7 +32,7 @@ pub mod verify;
|
|||
|
||||
pub use bind::{block_init_words, day_bytes, pow256_from_lane, target64_from_le256};
|
||||
pub use accept::{check as accept_program, AcceptReport, Reject};
|
||||
pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS};
|
||||
pub use memhard::{Cache, MemhardCpu, MixParams};
|
||||
pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS};
|
||||
pub use memhard::{cache_log2_words, dataset_log2_words, days_since_genesis, growth_doublings, Cache, MemhardCpu, MixParams, Shape};
|
||||
pub use seed::{fnv1a64, seed_words, SplitMix64};
|
||||
pub use verify::{hash_warp, interpret_warp_init, verify_block, DatasetMode, DatasetSource, Epoch};
|
||||
|
|
|
|||
|
|
@ -13,7 +13,7 @@
|
|||
|
||||
use igneum_pow::emit::export_pack;
|
||||
use igneum_pow::generator::LoadClass;
|
||||
use igneum_pow::memhard::Cache;
|
||||
use igneum_pow::memhard::{Cache, Shape};
|
||||
use igneum_pow::seed::day_key;
|
||||
use igneum_pow::verify::{DatasetMode, Epoch, DEFAULT_DATASET_LOG2};
|
||||
use std::time::Instant;
|
||||
|
|
@ -31,6 +31,8 @@ struct Args {
|
|||
epoch_hex: Option<String>,
|
||||
day_hex: Option<String>,
|
||||
class: LoadClass,
|
||||
/// Days since genesis for the cache growth rule of a class with `growth` (0: the genesis cache).
|
||||
days: u64,
|
||||
}
|
||||
|
||||
fn usage() -> ! {
|
||||
|
|
@ -42,7 +44,8 @@ fn usage() -> ! {
|
|||
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
|
||||
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
|
||||
\x20 show the accepted program, one instruction per line\n\
|
||||
\x20 --class C load class (read-width experiment): v2 (default), w4, w16, w64, w64x4, or p4,p16,p64[xN]"
|
||||
\x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
|
||||
\x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)"
|
||||
);
|
||||
std::process::exit(2)
|
||||
}
|
||||
|
|
@ -61,6 +64,7 @@ fn parse() -> Args {
|
|||
epoch_hex: None,
|
||||
day_hex: None,
|
||||
class: LoadClass::V2,
|
||||
days: 0,
|
||||
};
|
||||
let mut it = std::env::args().skip(1);
|
||||
a.cmd = it.next().unwrap_or_else(|| usage());
|
||||
|
|
@ -78,6 +82,7 @@ fn parse() -> Args {
|
|||
"--epoch-hex" => a.epoch_hex = Some(val()),
|
||||
"--day-hex" => a.day_hex = Some(val()),
|
||||
"--class" => a.class = LoadClass::parse(&val()).unwrap_or_else(|| usage()),
|
||||
"--days" => a.days = val().parse().unwrap_or_else(|_| usage()),
|
||||
_ => usage(),
|
||||
}
|
||||
}
|
||||
|
|
@ -93,7 +98,7 @@ fn main() {
|
|||
"accept" => accept(&a),
|
||||
"show" => show(&a),
|
||||
"hash" => {
|
||||
let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class);
|
||||
let e = Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days);
|
||||
println!("{:016x}", e.hash(a.nonce as u32));
|
||||
}
|
||||
"hash-bound" => {
|
||||
|
|
@ -106,7 +111,7 @@ fn main() {
|
|||
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
|
||||
Epoch::from_seed_bytes_class(&eb, &db, "cli", a.class)
|
||||
}
|
||||
_ => Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class),
|
||||
_ => Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days),
|
||||
};
|
||||
let init = igneum_pow::bind::block_init_words(&prehash, a.nonce);
|
||||
println!("init words {}", init.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" "));
|
||||
|
|
@ -124,24 +129,34 @@ fn bench(a: &Args, mode: DatasetMode) {
|
|||
a.dataset_log2,
|
||||
mode.name()
|
||||
);
|
||||
let shape = Shape::for_class_day(&a.class, a.days);
|
||||
if mode == DatasetMode::MemoryHard {
|
||||
// Time the cache fill on its own first (one core), then build the epoch (which fills it again).
|
||||
let t0 = Instant::now();
|
||||
let c = Cache::fill(day_key(&a.day));
|
||||
let c = Cache::fill_log2(day_key(&a.day), shape.cache_log2_words);
|
||||
let fill_ms = t0.elapsed().as_secs_f64() * 1e3;
|
||||
println!("cache: fill {fill_ms:.1} ms on one core (2^26 words, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}", c.fnv1a64());
|
||||
println!(
|
||||
"cache: fill {fill_ms:.1} ms on one core (2^{} words, {} MiB, {} chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}",
|
||||
shape.cache_log2_words,
|
||||
shape.cache_words() * 4 / (1 << 20),
|
||||
c.segments(),
|
||||
c.fnv1a64()
|
||||
);
|
||||
drop(c);
|
||||
}
|
||||
let t0 = Instant::now();
|
||||
let e = Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class);
|
||||
let e = Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days);
|
||||
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
|
||||
println!(
|
||||
"program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, op mix {}; epoch built in {build_ms:.1} ms",
|
||||
"program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, mixer x{} ({} mixers/item), cache 2^{} words, op mix {}; epoch built in {build_ms:.1} ms",
|
||||
e.program.class.name(),
|
||||
e.program.loads_per_hash(),
|
||||
e.program.bytes_per_hash(),
|
||||
e.program.width_counts(),
|
||||
e.program.items_per_warp(),
|
||||
shape.mixer_mult,
|
||||
shape.mixers_per_item(),
|
||||
shape.cache_log2_words,
|
||||
e.program.op_mix()
|
||||
);
|
||||
let bases = [0u32, 4096, 1_000_000];
|
||||
|
|
@ -175,7 +190,7 @@ fn export(a: &Args, mode: DatasetMode) {
|
|||
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
|
||||
(Epoch::from_seed_bytes_class(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}"), a.class), format!("bytes:{dh}"))
|
||||
}
|
||||
_ => (Epoch::new_class(&a.seed, &a.day, mode, a.dataset_log2, a.class), a.day.clone()),
|
||||
_ => (Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days), a.day.clone()),
|
||||
};
|
||||
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
|
||||
println!("igneum-pow export {out}");
|
||||
|
|
|
|||
|
|
@ -3,12 +3,18 @@
|
|||
//! ARX-multiply mixer. The verifier holds the cache and never the dataset.
|
||||
//!
|
||||
//! All arithmetic is on u32 modulo 2^32. Rotations are by 1..31 at every call site.
|
||||
//!
|
||||
//! Counter ASIC 2.0 (5 October 2026, `docs/plans/mixer-x4.md`, behind the program class): the construction has a
|
||||
//! [`Shape`], the mixer multiplier `m` and the cache size. Under `m` every mixer application of an item becomes
|
||||
//! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the
|
||||
//! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit.
|
||||
|
||||
use crate::generator::LoadClass;
|
||||
use crate::seed::{day_key, fnv1a64_words, SplitMix64};
|
||||
|
||||
pub const CACHE_LOG2_WORDS: usize = 26;
|
||||
pub const CACHE_SEGMENT_LOG2_LINES: usize = 6;
|
||||
/// 2^26 words = 256 MiB.
|
||||
/// 2^26 words = 256 MiB (the version 2 cache, and the v3 cache until the first dataset doubling).
|
||||
pub const CACHE_WORDS: usize = 1 << CACHE_LOG2_WORDS;
|
||||
/// 2^22 lines of 16 words.
|
||||
pub const CACHE_LINES: usize = CACHE_WORDS >> 4;
|
||||
|
|
@ -24,6 +30,94 @@ pub const CHACHA_SIGMA: [u32; 4] = [0x61707865, 0x3320646e, 0x79622d32, 0x6b2065
|
|||
/// "Igne", "umMH".
|
||||
pub const CACHE_TAG: [u32; 2] = [0x49676e65, 0x756d4d48];
|
||||
|
||||
/// The shape of the item derivation and of the cache: the mixer multiplier and the cache size.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
|
||||
pub struct Shape {
|
||||
/// Mixer applications per round (and after the last read): 1 under version 2, 4 under class v3.
|
||||
pub mixer_mult: u32,
|
||||
/// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3).
|
||||
pub cache_log2_words: u32,
|
||||
}
|
||||
|
||||
impl Shape {
|
||||
/// Version 2: one mixer application per round, a 2^26-word cache.
|
||||
pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32 };
|
||||
|
||||
/// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule).
|
||||
pub fn for_class(class: &LoadClass) -> Shape {
|
||||
Shape::for_class_day(class, 0)
|
||||
}
|
||||
|
||||
/// The shape of a load class on day `days_since_genesis` of the chain: the class's multiplier, and the cache
|
||||
/// of [`cache_log2_words`] when the class has the growth rule, else 2^26 words.
|
||||
pub fn for_class_day(class: &LoadClass, days_since_genesis: u64) -> Shape {
|
||||
Shape {
|
||||
mixer_mult: class.mixer_mult(),
|
||||
cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 },
|
||||
}
|
||||
}
|
||||
|
||||
pub fn is_v2(&self) -> bool {
|
||||
*self == Shape::V2
|
||||
}
|
||||
pub fn cache_words(&self) -> usize {
|
||||
1usize << self.cache_log2_words
|
||||
}
|
||||
pub fn cache_lines(&self) -> usize {
|
||||
self.cache_words() >> 4
|
||||
}
|
||||
pub fn cache_line_mask(&self) -> u32 {
|
||||
(self.cache_lines() - 1) as u32
|
||||
}
|
||||
pub fn cache_segments(&self) -> usize {
|
||||
self.cache_lines() >> CACHE_SEGMENT_LOG2_LINES
|
||||
}
|
||||
pub fn log2_segments(&self) -> u32 {
|
||||
self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32
|
||||
}
|
||||
/// Mixer applications per item: `(ITEM_ROUNDS + 1) x m`.
|
||||
pub fn mixers_per_item(&self) -> u32 {
|
||||
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult
|
||||
}
|
||||
}
|
||||
|
||||
// --------------------------------------------------------------------------------------------------------------
|
||||
// Dataset growth, option C (spec 01 section 1.13.3 option (b) with the cache tied to the dataset's doublings)
|
||||
// --------------------------------------------------------------------------------------------------------------
|
||||
|
||||
/// Days per year of the growth schedule: one year = 31,536,000 DAA seconds of 86,400 (spec 01 section 1.13.3).
|
||||
pub const GROWTH_DAYS_PER_YEAR: u64 = 365;
|
||||
/// The linear schedule of 1.13.3, 2 GiB at genesis plus 0.5 GiB per year, is `G x (1 + d / 1460)` for the genesis
|
||||
/// size `G` and the day `d`: it doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at
|
||||
/// day 10,220 (year 28) and 16x at day 21,900 (year 60).
|
||||
pub const GROWTH_DOUBLING_DAYS: u64 = 4 * GROWTH_DAYS_PER_YEAR;
|
||||
|
||||
/// The number of dataset doublings reached by day `days_since_genesis` of the chain: `floor(log2(1 + d / 1460))`,
|
||||
/// in integers (`1 + d / 1460` rounded down, then its integer log2, which equals the real log2's floor because a
|
||||
/// power of two is an integer). 0 until day 1,459; 1 from day 1,460 (year 4); 2 from day 4,380 (year 12).
|
||||
pub fn growth_doublings(days_since_genesis: u64) -> u32 {
|
||||
(1 + days_since_genesis / GROWTH_DOUBLING_DAYS).ilog2()
|
||||
}
|
||||
|
||||
/// The cache size on day `d` under option C: 2^26 words doubled once per dataset doubling (256 MiB, 512 MiB from
|
||||
/// year 4, 1 GiB from year 12).
|
||||
pub fn cache_log2_words(days_since_genesis: u64) -> u32 {
|
||||
CACHE_LOG2_WORDS as u32 + growth_doublings(days_since_genesis)
|
||||
}
|
||||
|
||||
/// The dataset size on day `d` under option (b) of 1.13.3: the genesis size (2^`genesis_log2_words` words: 28 for
|
||||
/// the 1 GiB packs and the devnet, 29 for the designed 2 GiB) doubled once per doubling of the linear schedule. The
|
||||
/// result is capped at 32 (the item index is 32 bits, spec 1.13.3).
|
||||
pub fn dataset_log2_words(genesis_log2_words: u32, days_since_genesis: u64) -> u32 {
|
||||
(genesis_log2_words + growth_doublings(days_since_genesis)).min(32)
|
||||
}
|
||||
|
||||
/// Days since genesis from two day indices of `bind::day_index` (the header's `timestamp_ms / 86,400,000`): the
|
||||
/// day of the block and the day of the genesis header. A block before the genesis day (clock skew) is day 0.
|
||||
pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 {
|
||||
day_index.saturating_sub(genesis_day_index)
|
||||
}
|
||||
|
||||
#[inline(always)]
|
||||
fn rotl(x: u32, n: u32) -> u32 {
|
||||
x.rotate_left(n)
|
||||
|
|
@ -66,17 +160,23 @@ pub fn chacha_block(x: &[u32; 16]) -> [u32; 16] {
|
|||
y
|
||||
}
|
||||
|
||||
/// Mixer parameters drawn from the day key. Draw order: ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15].
|
||||
/// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order:
|
||||
/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct MixParams {
|
||||
pub key: [u32; 8],
|
||||
pub rot: [u32; 8],
|
||||
pub mul: [u32; 16],
|
||||
pub rc: [u32; 16],
|
||||
pub shape: Shape,
|
||||
}
|
||||
|
||||
impl MixParams {
|
||||
/// Version 2 shape.
|
||||
pub fn new(key: [u32; 8]) -> Self {
|
||||
Self::with_shape(key, Shape::V2)
|
||||
}
|
||||
pub fn with_shape(key: [u32; 8], shape: Shape) -> Self {
|
||||
let mut rng = SplitMix64::new(key[0] as u64 | ((key[1] as u64) << 32));
|
||||
let mut rot = [0u32; 8];
|
||||
let mut mul = [0u32; 16];
|
||||
|
|
@ -90,7 +190,7 @@ impl MixParams {
|
|||
for c in rc.iter_mut() {
|
||||
*c = rng.next() as u32;
|
||||
}
|
||||
Self { key, rot, mul, rc }
|
||||
Self { key, rot, mul, rc, shape }
|
||||
}
|
||||
/// Parameters for a day string: the key is `seed_words("day/" + day)`.
|
||||
pub fn for_day(day: &str) -> Self {
|
||||
|
|
@ -104,6 +204,13 @@ pub fn round_key(r: usize) -> u32 {
|
|||
((r + 1) as u32).wrapping_mul(0x9E3779B9)
|
||||
}
|
||||
|
||||
/// The round key of application `j` (0 <= j < m) of round `r` under multiplier `m`: `round_key(r * m + j)`. For
|
||||
/// `m = 1` this is `round_key(r)`, version 2's key.
|
||||
#[inline(always)]
|
||||
pub fn round_key_mult(r: usize, j: usize, m: usize) -> u32 {
|
||||
round_key(r * m + j)
|
||||
}
|
||||
|
||||
/// `M_r` on 16 words in place: per word `(s ^ (RC + rk)) * MUL`, then one ChaCha-shaped double round with
|
||||
/// the four column rotations `ROT[0..3]` and the four diagonal rotations `ROT[4..7]`.
|
||||
#[inline(always)]
|
||||
|
|
@ -122,9 +229,11 @@ pub fn mixer(s: &mut [u32; 16], rk: u32, mp: &MixParams) {
|
|||
qr(s, 3, 4, 9, 14, r[4], r[5], r[6], r[7]);
|
||||
}
|
||||
|
||||
/// The 256 MiB cache for one day key.
|
||||
/// The cache for one day key: 2^log2_words words (256 MiB under version 2).
|
||||
pub struct Cache {
|
||||
pub key: [u32; 8],
|
||||
pub log2_words: u32,
|
||||
line_mask: u32,
|
||||
words: Vec<u32>,
|
||||
}
|
||||
|
||||
|
|
@ -152,13 +261,22 @@ impl Cache {
|
|||
}
|
||||
}
|
||||
|
||||
/// The whole cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order.
|
||||
/// The version 2 cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order.
|
||||
pub fn fill(key: [u32; 8]) -> Cache {
|
||||
let mut words = vec![0u32; CACHE_WORDS];
|
||||
for seg in 0..CACHE_SEGMENTS {
|
||||
Self::fill_log2(key, CACHE_LOG2_WORDS as u32)
|
||||
}
|
||||
|
||||
/// A cache of 2^`log2_words` words (26, 27 or 28 under the growth rule; smaller sizes for tests): 2^(log2 - 10)
|
||||
/// independent chains of 64 lines, the same chain function at every size, so a larger cache's first segments
|
||||
/// are the smaller cache's segments word for word.
|
||||
pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache {
|
||||
assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30");
|
||||
let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words };
|
||||
let mut words = vec![0u32; shape.cache_words()];
|
||||
for seg in 0..shape.cache_segments() {
|
||||
Self::fill_segment(&mut words, seg, &key);
|
||||
}
|
||||
Cache { key, words }
|
||||
Cache { key, log2_words, line_mask: shape.cache_line_mask(), words }
|
||||
}
|
||||
|
||||
pub fn for_day(day: &str) -> Cache {
|
||||
|
|
@ -170,10 +288,20 @@ impl Cache {
|
|||
&self.words
|
||||
}
|
||||
|
||||
/// Cache line `a` (0 <= a < 2^22) as 16 words.
|
||||
pub fn lines(&self) -> usize {
|
||||
self.words.len() >> 4
|
||||
}
|
||||
pub fn line_mask(&self) -> u32 {
|
||||
self.line_mask
|
||||
}
|
||||
pub fn segments(&self) -> usize {
|
||||
self.lines() >> CACHE_SEGMENT_LOG2_LINES
|
||||
}
|
||||
|
||||
/// Cache line `a` (masked to the cache's lines) as 16 words.
|
||||
#[inline(always)]
|
||||
pub fn line(&self, a: u32) -> &[u32] {
|
||||
let o = (a & CACHE_LINE_MASK) as usize * 16;
|
||||
let o = (a & self.line_mask) as usize * 16;
|
||||
&self.words[o..o + 16]
|
||||
}
|
||||
|
||||
|
|
@ -184,10 +312,13 @@ impl Cache {
|
|||
}
|
||||
|
||||
/// Derive `ts.len()` items into `out`, all chains interleaved round by round so the cache-line misses of
|
||||
/// independent items overlap in the memory system (`deriveItems` in the Swift).
|
||||
/// independent items overlap in the memory system (`deriveItems` in the Swift). Under multiplier `m`
|
||||
/// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its
|
||||
/// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2.
|
||||
pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) {
|
||||
let n = ts.len();
|
||||
debug_assert!(out.len() >= n);
|
||||
let m = mp.shape.mixer_mult as usize;
|
||||
for k in 0..n {
|
||||
let s = &mut out[k];
|
||||
let t = ts[k];
|
||||
|
|
@ -197,9 +328,11 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32;
|
|||
}
|
||||
}
|
||||
for r in 0..ITEM_ROUNDS {
|
||||
let rk = round_key(r);
|
||||
for s in out[..n].iter_mut() {
|
||||
mixer(s, rk, mp);
|
||||
for j in 0..m {
|
||||
let rk = round_key_mult(r, j, m);
|
||||
for s in out[..n].iter_mut() {
|
||||
mixer(s, rk, mp);
|
||||
}
|
||||
}
|
||||
for s in out[..n].iter_mut() {
|
||||
let line = cache.line(s[0]);
|
||||
|
|
@ -208,9 +341,11 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32;
|
|||
}
|
||||
}
|
||||
}
|
||||
let rk = round_key(ITEM_ROUNDS);
|
||||
for s in out[..n].iter_mut() {
|
||||
mixer(s, rk, mp);
|
||||
for j in 0..m {
|
||||
let rk = round_key_mult(ITEM_ROUNDS, j, m);
|
||||
for s in out[..n].iter_mut() {
|
||||
mixer(s, rk, mp);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -221,7 +356,7 @@ pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
|
|||
out[0]
|
||||
}
|
||||
|
||||
/// The CPU verifier's view of the memory-hard dataset: the mixer parameters and the 256 MiB cache.
|
||||
/// The CPU verifier's view of the memory-hard dataset: the mixer parameters (with the shape) and the cache.
|
||||
pub struct MemhardCpu {
|
||||
pub params: MixParams,
|
||||
pub cache: Cache,
|
||||
|
|
@ -231,12 +366,19 @@ pub struct MemhardCpu {
|
|||
pub const FETCH_MAX: usize = 64;
|
||||
|
||||
impl MemhardCpu {
|
||||
/// Version 2 shape.
|
||||
pub fn new(key: [u32; 8]) -> Self {
|
||||
Self { params: MixParams::new(key), cache: Cache::fill(key) }
|
||||
Self::with_shape(key, Shape::V2)
|
||||
}
|
||||
pub fn with_shape(key: [u32; 8], shape: Shape) -> Self {
|
||||
Self { params: MixParams::with_shape(key, shape), cache: Cache::fill_log2(key, shape.cache_log2_words) }
|
||||
}
|
||||
pub fn for_day(day: &str) -> Self {
|
||||
Self::new(day_key(day))
|
||||
}
|
||||
pub fn shape(&self) -> Shape {
|
||||
self.params.shape
|
||||
}
|
||||
/// `dataset[w] = item(w >> 4)[w & 15]`.
|
||||
pub fn word(&self, w: u32) -> u32 {
|
||||
derive_item(w >> 4, &self.params, &self.cache)[(w & 15) as usize]
|
||||
|
|
@ -313,6 +455,7 @@ mod tests {
|
|||
assert_eq!(mp.rc[0], 0xbab68293);
|
||||
assert_eq!(mp.rc[15], 0x31b49ee2);
|
||||
assert!(mp.mul.iter().all(|m| m & 1 == 1));
|
||||
assert_eq!(mp.shape, Shape::V2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
|
@ -338,4 +481,96 @@ mod tests {
|
|||
]
|
||||
);
|
||||
}
|
||||
|
||||
/// Option C: the schedule table of `docs/plans/mixer-x4.md` (day -> doublings, cache words, dataset words at a
|
||||
/// 2^28 genesis). The doublings fall at years 4 and 12 exactly, never a day early.
|
||||
#[test]
|
||||
fn growth_schedule_table() {
|
||||
let table: [(u64, u32, u32, u32); 12] = [
|
||||
(0, 0, 26, 28),
|
||||
(1, 0, 26, 28),
|
||||
(365, 0, 26, 28),
|
||||
(1_459, 0, 26, 28),
|
||||
(1_460, 1, 27, 29),
|
||||
(2_920, 1, 27, 29),
|
||||
(4_379, 1, 27, 29),
|
||||
(4_380, 2, 28, 30),
|
||||
(10_219, 2, 28, 30),
|
||||
(10_220, 3, 29, 31),
|
||||
(21_900, 4, 30, 32),
|
||||
(100_000, 6, 32, 32),
|
||||
];
|
||||
for (d, k, c, s) in table {
|
||||
assert_eq!(growth_doublings(d), k, "day {d}");
|
||||
assert_eq!(cache_log2_words(d), c, "day {d}");
|
||||
assert_eq!(dataset_log2_words(28, d), s, "day {d}");
|
||||
}
|
||||
// the designed 2 GiB genesis: 2^29 words, 2^30 at year 4, 2^31 at year 12
|
||||
assert_eq!(dataset_log2_words(29, 0), 29);
|
||||
assert_eq!(dataset_log2_words(29, 1_460), 30);
|
||||
assert_eq!(dataset_log2_words(29, 4_380), 31);
|
||||
// the linear schedule itself: 2 GiB x (1 + d / 1460) crosses 4 GiB at day 1,460 and 8 GiB at day 4,380
|
||||
for d in [1_459u64, 1_460, 4_379, 4_380] {
|
||||
let bytes = 2u64 * (1 << 30) + (1u64 << 29) * d / 365;
|
||||
let k = (bytes / (2u64 << 30)).ilog2();
|
||||
assert_eq!(growth_doublings(d), k, "day {d}: linear {bytes} bytes");
|
||||
}
|
||||
assert_eq!(days_since_genesis(20_730, 20_729), 1);
|
||||
assert_eq!(days_since_genesis(20_729, 20_729), 0);
|
||||
assert_eq!(days_since_genesis(20_000, 20_729), 0);
|
||||
let v2 = Shape::for_class_day(&LoadClass::V2, 100_000);
|
||||
assert_eq!(v2, Shape::V2);
|
||||
let v3 = Shape::for_class_day(&LoadClass::MX4, 0);
|
||||
assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26 });
|
||||
assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27);
|
||||
assert_eq!(v3.mixers_per_item(), 36);
|
||||
assert_eq!(Shape::V2.mixers_per_item(), 9);
|
||||
assert_eq!(Shape::V2.cache_segments(), CACHE_SEGMENTS);
|
||||
assert_eq!(Shape::V2.cache_line_mask(), CACHE_LINE_MASK);
|
||||
assert_eq!(Shape::V2.log2_segments(), 16);
|
||||
}
|
||||
|
||||
/// The multiplied mixer, restated by hand on a small cache: `m` applications with keys `round_key(r m + j)`
|
||||
/// before every read, the same 8 reads; `m = 1` is `derive_item` of version 2 word for word; a larger cache's
|
||||
/// first segments equal the smaller cache's.
|
||||
#[test]
|
||||
fn mixer_mult_by_hand() {
|
||||
let key = day_key("2026-10-03");
|
||||
let small = Cache::fill_log2(key, 16);
|
||||
let big = Cache::fill_log2(key, 18);
|
||||
assert_eq!(&big.words()[..small.words().len()], small.words());
|
||||
assert_eq!(small.segments(), 64);
|
||||
assert_eq!(small.line_mask(), 4095);
|
||||
for m in [1u32, 2, 4] {
|
||||
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16 });
|
||||
for t in [0u32, 1, 12_345, u32::MAX] {
|
||||
let got = derive_item(t, &mp, &small);
|
||||
let mut s = [0u32; 16];
|
||||
s[..8].copy_from_slice(&key);
|
||||
for i in 0..8 {
|
||||
s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
|
||||
}
|
||||
for r in 0..8usize {
|
||||
for j in 0..m as usize {
|
||||
mixer(&mut s, round_key(r * m as usize + j), &mp);
|
||||
}
|
||||
let line = small.line(s[0]);
|
||||
for i in 0..16 {
|
||||
s[i] ^= line[i];
|
||||
}
|
||||
}
|
||||
for j in 0..m as usize {
|
||||
mixer(&mut s, round_key(8 * m as usize + j), &mp);
|
||||
}
|
||||
assert_eq!(got, s, "m {m} t {t}");
|
||||
}
|
||||
}
|
||||
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16 });
|
||||
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16 });
|
||||
assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small));
|
||||
assert_eq!(round_key_mult(0, 0, 1), round_key(0));
|
||||
assert_eq!(round_key_mult(8, 0, 1), round_key(8));
|
||||
assert_eq!(round_key_mult(2, 3, 4), round_key(11));
|
||||
assert_eq!(round_key_mult(8, 3, 4), round_key(35));
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
//! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs).
|
||||
|
||||
use crate::generator::{generate, generate_class, Instr, LoadClass, Op, Program, ProgramClass, ITERATIONS, LANES};
|
||||
use crate::memhard::MemhardCpu;
|
||||
use crate::memhard::{MemhardCpu, Shape};
|
||||
use crate::seed::day_key;
|
||||
|
||||
/// Read-width experiment (5 October 2026): a `load` of `W` words folds every word into `dst`:
|
||||
|
|
@ -150,21 +150,37 @@ pub struct DatasetSource {
|
|||
impl DatasetSource {
|
||||
/// Build the source for a day. Memory-hard mode fills the 256 MiB cache on the calling thread.
|
||||
pub fn new(day: &str, mode: DatasetMode, log2_words: u32) -> Self {
|
||||
let mut ds = Self::from_key(day_key(day), mode, log2_words);
|
||||
Self::new_shape(day, mode, log2_words, Shape::V2)
|
||||
}
|
||||
|
||||
/// [`DatasetSource::new`] with the construction's shape (mixer multiplier, cache size; Counter ASIC 2.0).
|
||||
pub fn new_shape(day: &str, mode: DatasetMode, log2_words: u32, shape: Shape) -> Self {
|
||||
let mut ds = Self::from_key_shape(day_key(day), mode, log2_words, shape);
|
||||
ds.key_bytes = format!("day/{day}").into_bytes();
|
||||
ds
|
||||
}
|
||||
|
||||
pub fn from_key(key: [u32; 8], mode: DatasetMode, log2_words: u32) -> Self {
|
||||
Self::from_key_shape(key, mode, log2_words, Shape::V2)
|
||||
}
|
||||
|
||||
/// [`DatasetSource::from_key`] with the construction's shape. Memory-hard mode fills a cache of
|
||||
/// `2^shape.cache_log2_words` words on the calling thread.
|
||||
pub fn from_key_shape(key: [u32; 8], mode: DatasetMode, log2_words: u32, shape: Shape) -> Self {
|
||||
assert!((4..=32).contains(&log2_words), "dataset log2 must be in 4..=32");
|
||||
let mask = if log2_words == 32 { u32::MAX } else { (1u32 << log2_words) - 1 };
|
||||
let dataset = match mode {
|
||||
DatasetMode::ClosedForm => Dataset::ClosedForm { d0: key[0], d1: key[1] },
|
||||
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::new(key)),
|
||||
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::with_shape(key, shape)),
|
||||
};
|
||||
Self { log2_words, mask, key, key_bytes: Vec::new(), dataset }
|
||||
}
|
||||
|
||||
/// The shape of the memory-hard construction ([`Shape::V2`] for the closed form, which has none).
|
||||
pub fn shape(&self) -> Shape {
|
||||
self.memhard().map(|m| m.shape()).unwrap_or(Shape::V2)
|
||||
}
|
||||
|
||||
pub fn mode(&self) -> DatasetMode {
|
||||
match self.dataset {
|
||||
Dataset::ClosedForm { .. } => DatasetMode::ClosedForm,
|
||||
|
|
@ -417,9 +433,10 @@ pub struct Epoch {
|
|||
/// Default dataset size: 2^28 words = 1 GiB.
|
||||
pub const DEFAULT_DATASET_LOG2: u32 = 28;
|
||||
|
||||
/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it).
|
||||
/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it). The node's
|
||||
/// entry; the same function as `memhard::days_since_genesis`.
|
||||
pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 {
|
||||
day_index.saturating_sub(genesis_day_index)
|
||||
crate::memhard::days_since_genesis(day_index, genesis_day_index)
|
||||
}
|
||||
|
||||
impl Epoch {
|
||||
|
|
@ -427,9 +444,17 @@ impl Epoch {
|
|||
Self { program: generate(seed), dataset: DatasetSource::new(day, mode, dataset_log2) }
|
||||
}
|
||||
|
||||
/// [`Epoch::new`] with a load class (read-width experiment).
|
||||
/// [`Epoch::new`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer multiplier
|
||||
/// shapes the dataset, the cache is the genesis size since a string day has no day index).
|
||||
pub fn new_class(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass) -> Self {
|
||||
Self { program: generate_class(seed, class), dataset: DatasetSource::new(day, mode, dataset_log2) }
|
||||
Self::new_class_day(seed, day, mode, dataset_log2, class, 0)
|
||||
}
|
||||
|
||||
/// [`Epoch::new_class`] on day `days_since_genesis` of the growth schedule (the cache of
|
||||
/// `memhard::cache_log2_words` for a class with the growth rule; the dataset size is the caller's).
|
||||
pub fn new_class_day(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass, days_since_genesis: u64) -> Self {
|
||||
let shape = Shape::for_class_day(&class, days_since_genesis);
|
||||
Self { program: generate_class(seed, class), dataset: DatasetSource::new_shape(day, mode, dataset_log2, shape) }
|
||||
}
|
||||
|
||||
/// The production shape: memory-hard, 1 GiB dataset.
|
||||
|
|
@ -445,11 +470,24 @@ impl Epoch {
|
|||
Self::from_seed_bytes_class(epoch_seed, day_bytes, label, LoadClass::V2)
|
||||
}
|
||||
|
||||
/// [`Epoch::from_seed_bytes`] with a load class (read-width experiment).
|
||||
/// [`Epoch::from_seed_bytes`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer
|
||||
/// multiplier shapes the dataset). Day 0 of the growth schedule: the 2^26-word cache and the 2^28-word dataset,
|
||||
/// which is every devnet pack and vector. A node past the first doubling calls [`Epoch::from_seed_bytes_day`].
|
||||
pub fn from_seed_bytes_class(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass) -> Self {
|
||||
Self::from_seed_bytes_day(epoch_seed, day_bytes, label, class, 0, DEFAULT_DATASET_LOG2)
|
||||
}
|
||||
|
||||
/// The chain's shape on day `days_since_genesis` (`memhard::days_since_genesis(day_index(header), day_index(genesis))`,
|
||||
/// the node's two day indices): the program of the class, and under the class's growth rule the cache of
|
||||
/// `memhard::cache_log2_words(d)` and the dataset of `memhard::dataset_log2_words(genesis_dataset_log2, d)`
|
||||
/// (the genesis size is 28 for the 1 GiB devnet, 29 for the designed 2 GiB). Without the growth rule the cache
|
||||
/// is 2^26 words and the dataset `2^genesis_dataset_log2` on every day.
|
||||
pub fn from_seed_bytes_day(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> Self {
|
||||
let program = crate::generator::generate_from_seed_bytes_class(label, epoch_seed, class);
|
||||
let key = crate::seed::seed_words_from_bytes(day_bytes);
|
||||
let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2);
|
||||
let shape = Shape::for_class_day(&class, days_since_genesis);
|
||||
let dataset_log2 = if class.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 };
|
||||
let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape);
|
||||
dataset.key_bytes = day_bytes.to_vec();
|
||||
Self { program, dataset }
|
||||
}
|
||||
|
|
@ -477,13 +515,18 @@ impl Epoch {
|
|||
|
||||
/// [`Epoch::chain_dataset`] with the day's position since genesis and the network's genesis dataset size: the
|
||||
/// entry the node's engine and the miner's export build every day cache through, so the cache growth schedule
|
||||
/// of spec 01 section 1.13.3 has one place to act. SEAM (ca2-node, 5 October 2026): the body here builds the
|
||||
/// genesis-size cache for every day; the ca2-mixer branch fills the growth rule (`growth_doublings`) and the
|
||||
/// class v3 item construction, keeping this signature.
|
||||
/// of spec 01 section 1.13.3 has one place to act (ca2-mixer, 5 October 2026, `docs/plans/mixer-x4.md`): the
|
||||
/// class's load class gives the mixer multiplier and whether the growth rule applies (`Shape::for_class_day`);
|
||||
/// under the rule the cache is `2^memhard::cache_log2_words(d)` words and the dataset
|
||||
/// `2^memhard::dataset_log2_words(genesis_dataset_log2, d)`; without it (class v2) the cache is 2^26 words and
|
||||
/// the dataset the genesis size on every day. `days_since_genesis` is [`days_since_genesis`] of the block's and
|
||||
/// the genesis header's day indices.
|
||||
pub fn chain_dataset_day(day_bytes: &[u8], class: ProgramClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> DatasetSource {
|
||||
let _ = (class, days_since_genesis);
|
||||
let lc = class.load_class();
|
||||
let shape = Shape::for_class_day(&lc, days_since_genesis);
|
||||
let dataset_log2 = if lc.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 };
|
||||
let key = crate::seed::seed_words_from_bytes(day_bytes);
|
||||
let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, genesis_dataset_log2);
|
||||
let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape);
|
||||
dataset.key_bytes = day_bytes.to_vec();
|
||||
dataset
|
||||
}
|
||||
|
|
|
|||
|
|
@ -31,6 +31,9 @@ typedef struct {
|
|||
char loadClass[64];
|
||||
char programClass[8]; /* IGNEUM_PROGRAM_CLASS: "v2" or "v3" (Counter ASIC 2.0); absent = the generator's class */
|
||||
char eraHex[65]; /* IGNEUM_ERA_SEED_HEX of a class v3 chain pack; empty otherwise */
|
||||
// Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation (IGNEUM_MIXER_MULT, 1 when absent:
|
||||
// version 2; 4 under class v3). The emitted memhard.h / kernel.cl carry it in their text; this is for the log lines.
|
||||
uint32_t mixerMult;
|
||||
// seeds.txt (or program.h): the seeds as the worker protocol carries them
|
||||
char epochHex[65];
|
||||
char dayHex[PF_HEX_CAP];
|
||||
|
|
@ -296,6 +299,7 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
|
|||
pk->persistent = 0; pf_define_u32(prog, "IGNEUM_PERSISTENT_WARPS", &pk->persistent);
|
||||
pk->scratchWordsPerLane = 8192; pf_define_u32(prog, "IGNEUM_SCRATCH_WORDS_PER_LANE", &pk->scratchWordsPerLane);
|
||||
strcpy(pk->loadClass, "v2"); pf_define_str(prog, "IGNEUM_LOAD_CLASS", pk->loadClass, sizeof(pk->loadClass));
|
||||
pk->mixerMult = 1; pf_define_u32(prog, "IGNEUM_MIXER_MULT", &pk->mixerMult);
|
||||
if (pf_define_words(prog, "IGNEUM_KEY_INIT", pk->keyw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_KEY_INIT with 8 words"); }
|
||||
if (!pf_define_str(prog, "IGNEUM_SEED_STRING", pk->seedString, sizeof(pk->seedString))) strncpy(pk->seedString, "(no IGNEUM_SEED_STRING)", sizeof(pk->seedString) - 1);
|
||||
pf_define_str(prog, "IGNEUM_SEED_BYTES_HEX", ehex, sizeof(ehex));
|
||||
|
|
|
|||
|
|
@ -61,6 +61,8 @@ let scratchOps = Int(defineU32("IGNEUM_SCRATCH_OPS") ?? 0)
|
|||
let persistent = (defineU32("IGNEUM_PERSISTENT_WARPS") ?? 0) == 1
|
||||
let scratchWordsPerLane = Int(defineU32("IGNEUM_SCRATCH_WORDS_PER_LANE") ?? 8192)
|
||||
let className = defineStr("IGNEUM_LOAD_CLASS") ?? "v2"
|
||||
// Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation, 1 when absent (version 2), 4 under class v3
|
||||
let mixerMult = Int(defineU32("IGNEUM_MIXER_MULT") ?? 1)
|
||||
let seedString = defineStr("IGNEUM_SEED_STRING") ?? "?"
|
||||
let programId = defineStr("IGNEUM_PROGRAM_ID") ?? ""
|
||||
|
||||
|
|
@ -200,7 +202,7 @@ for b in 0..<opts.batches {
|
|||
let hashes = Double(nonces) * Double(opts.batches)
|
||||
let mhsWall = hashes / wallSum / 1e3, mhsGpu = hashes / gpuSum / 1e3
|
||||
let packName = (opts.pack as NSString).lastPathComponent
|
||||
print("pack \(packName) seed \"\(seedString)\" id \(programId) class \(className): loads/hash \(loadsPerHash), dataset bytes/hash \(bytesPerHash), scratch ops/hash \(scratchOps * 8)")
|
||||
print("pack \(packName) seed \"\(seedString)\" id \(programId) class \(className): loads/hash \(loadsPerHash), dataset bytes/hash \(bytesPerHash), scratch ops/hash \(scratchOps * 8), mixer x\(mixerMult), cache 2^\(cacheLog2) words")
|
||||
print("device \(device.name); compile \(String(format: "%.0f", compileMs)) ms; cache fill \(String(format: "%.1f", cacheGpu)) ms GPU (\(String(format: "%.1f", cacheWall)) wall); dataset build \(String(format: "%.1f", buildGpu)) ms GPU (\(String(format: "%.1f", buildWall)) wall)")
|
||||
print("cache FNV-1a 64 \(String(format: "%016llx", cacheFnv)) \(cacheOk ? "PASS" : "FAIL"); dataset head and last \(dsOk ? "PASS" : "FAIL"); vectors standalone \(vecPass)/\(vecBases.count), in batch \(batchVecPass)/\(batchVecN)")
|
||||
print("warm-up batch \(nonces) hashes: \(String(format: "%.1f", warmGpu)) ms GPU, \(String(format: "%.1f", warmWall)) ms wall")
|
||||
|
|
|
|||
|
|
@ -2092,7 +2092,11 @@ int main(int argc, char** argv) {
|
|||
printf("day \"%s\" (d0 0x%08x, d1 0x%08x), pack dataset 2^%d words = %d MiB\n",
|
||||
IGNEUM_DAY_STRING, IGNEUM_DAY0, IGNEUM_DAY1, IGNEUM_DATASET_LOG2, packMib());
|
||||
#if IGNEUM_DATASET_MODE == 1
|
||||
printf("dataset construction: memory-hard (256 MiB ChaCha cache, %d dependent cache reads per 64-byte item; proto-metal/MEMHARD.md)\n", IGNEUM_ITEM_ROUNDS);
|
||||
#ifndef IGNEUM_MIXER_MULT
|
||||
#define IGNEUM_MIXER_MULT 1
|
||||
#endif
|
||||
printf("dataset construction: memory-hard (%u MiB ChaCha cache, %d dependent cache reads per 64-byte item, mixer x%d; proto-metal/MEMHARD.md)\n",
|
||||
(unsigned)(((uint64_t)CACHE_WORDS_HOST * 4u) >> 20), IGNEUM_ITEM_ROUNDS, IGNEUM_MIXER_MULT);
|
||||
cachePass = setupCache(&dv, di);
|
||||
#else
|
||||
printf("dataset construction: closed-form ds_elem (the original prototype dataset, not memory-hard)\n");
|
||||
|
|
|
|||
63
relay/playbooks/mixer-x4-5090-bench.ps1
Normal file
63
relay/playbooks/mixer-x4-5090-bench.ps1
Normal file
|
|
@ -0,0 +1,63 @@
|
|||
# Igneum run job (bench only): the class v3 dataset construction (mixer x4, docs/plans/mixer-x4.md) on PC 2's RTX 5090 (machine 1ccfe586), 5 October 2026.
|
||||
# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches
|
||||
# off ONLY the NVIDIA card in the app through POST <app.url>api/cards, waits for its worker to stop, runs the fetched
|
||||
# igneum-worker-cuda.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the
|
||||
# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed),
|
||||
# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs <job id>` shows them.
|
||||
# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus
|
||||
# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run).
|
||||
$ErrorActionPreference = 'Continue'
|
||||
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
|
||||
$jobs = Split-Path $env:IGNEUM_JOB_DIR
|
||||
$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005'
|
||||
$exe = Join-Path $fetched 'igneum-worker-cuda.exe'
|
||||
$packs = Join-Path $fetched 'packs-ca2-mixer'
|
||||
if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 }
|
||||
if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 }
|
||||
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
|
||||
if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 }
|
||||
Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force
|
||||
Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst"
|
||||
|
||||
# the app: switch off the NVIDIA card only, remember its settings
|
||||
$appDir = $env:IGNEUM_APP_DIR
|
||||
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
|
||||
$urlFile = Join-Path $appDir 'app.url'
|
||||
$url = $null
|
||||
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
|
||||
$card = $null
|
||||
if ($url) {
|
||||
try {
|
||||
$st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10
|
||||
$card = $st.mining.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1
|
||||
if (-not $card) { $card = $st.cards | Where-Object { $_.vendor -eq 'nvidia' } | Select-Object -First 1 }
|
||||
} catch { Say ("api/state: " + $_.Exception.Message) }
|
||||
}
|
||||
if ($card) {
|
||||
Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state)
|
||||
$body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) }
|
||||
$t = 0
|
||||
while ($t -lt 90) {
|
||||
Start-Sleep -Seconds 5; $t += 5
|
||||
try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { }
|
||||
}
|
||||
Write-Output ("RESULT card-off after " + $t + " s")
|
||||
Start-Sleep -Seconds 5
|
||||
} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
|
||||
|
||||
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
|
||||
# the probe ran in run-readwidth-5090-20261005 (bench-log, 5 October 2026); this second run is the bench only
|
||||
foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) {
|
||||
$d = Join-Path $packs $pk
|
||||
Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)"
|
||||
& $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" }
|
||||
& $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" }
|
||||
}
|
||||
& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
|
||||
|
||||
if ($card) {
|
||||
$body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) }
|
||||
}
|
||||
exit 0
|
||||
64
relay/playbooks/mixer-x4-9070-bench.ps1
Normal file
64
relay/playbooks/mixer-x4-9070-bench.ps1
Normal file
|
|
@ -0,0 +1,64 @@
|
|||
# Igneum run job (bench only): the class v3 dataset construction (mixer x4, docs/plans/mixer-x4.md) on PC 1's RX 9070 XT on the eGPU (machine ae432dc7), 5 October 2026.
|
||||
# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining; this script switches
|
||||
# off ONLY the NVIDIA card in the app through POST <app.url>api/cards, waits for its worker to stop, runs the fetched
|
||||
# igneum-worker-opencl.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the
|
||||
# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed),
|
||||
# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs <job id>` shows them.
|
||||
# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus
|
||||
# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run).
|
||||
$ErrorActionPreference = 'Continue'
|
||||
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
|
||||
$jobs = Split-Path $env:IGNEUM_JOB_DIR
|
||||
$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005'
|
||||
$exe = Join-Path $fetched 'igneum-worker-opencl.exe'
|
||||
$packs = Join-Path $fetched 'packs-ca2-mixer'
|
||||
if (-not (Test-Path $exe)) { Write-Output "RESULT error worker missing at $exe (the fetch job runs first)"; exit 2 }
|
||||
if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 }
|
||||
Write-Output "RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())"
|
||||
# the card's OpenCL device index on the current (3683.0) platform, from the worker's own list (the older platform's duplicate is marked dup)
|
||||
$list = & $exe --list 2>&1
|
||||
$list | ForEach-Object { "RESULT list $_" }
|
||||
$dev = $null
|
||||
foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; break } }
|
||||
if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 device in --list'; exit 2 }
|
||||
Write-Output "RESULT device $dev"
|
||||
|
||||
# the app: switch off the NVIDIA card only, remember its settings
|
||||
$appDir = $env:IGNEUM_APP_DIR
|
||||
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
|
||||
$urlFile = Join-Path $appDir 'app.url'
|
||||
$url = $null
|
||||
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
|
||||
$card = $null
|
||||
if ($url) {
|
||||
try {
|
||||
$st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10
|
||||
$card = $st.mining.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1
|
||||
if (-not $card) { $card = $st.cards | Where-Object { ($_.vendor -eq 'amd' -and $_.key -match 'gfx1201') } | Select-Object -First 1 }
|
||||
} catch { Say ("api/state: " + $_.Exception.Message) }
|
||||
}
|
||||
if ($card) {
|
||||
Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state)
|
||||
$body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) }
|
||||
$t = 0
|
||||
while ($t -lt 90) {
|
||||
Start-Sleep -Seconds 5; $t += 5
|
||||
try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { }
|
||||
}
|
||||
Write-Output ("RESULT card-off after " + $t + " s")
|
||||
Start-Sleep -Seconds 5
|
||||
} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
|
||||
|
||||
# the probe ran in run-readwidth-9070-20261005 (bench-log, 5 October 2026); this second run is the bench only
|
||||
foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) {
|
||||
$d = Join-Path $packs $pk
|
||||
Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)"
|
||||
& $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" }
|
||||
}
|
||||
|
||||
if ($card) {
|
||||
$body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
|
||||
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore: " + $_.Exception.Message) }
|
||||
}
|
||||
exit 0
|
||||
Loading…
Reference in a new issue