Merge branch 'ca2-soundness' into ca2-coord
# Conflicts: # docs/bench-log.md # proto-metal/packbench.swift
This commit is contained in:
commit
271cd63116
4 changed files with 1237 additions and 2 deletions
408
docs/analysis/scratch-soundness.md
Normal file
408
docs/analysis/scratch-soundness.md
Normal file
|
|
@ -0,0 +1,408 @@
|
|||
# Layer 3 soundness: the per-warp scratch with read-modify-writes
|
||||
|
||||
5 October 2026 (night), cryptographer role, Counter ASIC 2.0 plan step 4 (`docs/plans/counter-asic-2.md`). Branch
|
||||
`ca2-soundness` on top of `readwidth` b970dda (the scratch as a class parameter, 32 or 128 KiB per warp). Tests:
|
||||
`igneum-pow/tests/scratch.rs`; Metal runs through `proto-metal/packbench` on the M5 Max; commands and counts in
|
||||
`docs/bench-log.md` (entry of the same date). Nothing here touches the lottery hash as shipped: variant 5 is behind
|
||||
`LoadClass::scratch(k, kb)` and is never emitted by generator version 2.
|
||||
|
||||
Every figure below is measured (machine, date, command named) or cited; "approximate" marks a figure from memory.
|
||||
|
||||
## 0. The five findings
|
||||
|
||||
| # | Question | Finding | Status |
|
||||
|---|---|---|---|
|
||||
| 1 | Is what is written uniform and beyond a chip's precomputation? | The fill is a bijection of the lane nonce, the rewrite a bijection of the fold value in each word; written words show no bit bias over 3 to 12 million rewrites per class (worst 3.63 sigma of 6). The fill IS precomputable, by design, and at 64 slots 78.5 percent of reads are fill reads. | sound as a function; see 2 for what that means |
|
||||
| 2 | Does any short cut avoid the writes? | No short cut inside a unit: a slot after d read-modify-writes needs all d fold values (replay test). But the live state is bounded by the read-modify-write count, not by the scratch size, because CPU verification resets the scratch per unit: 64 to 320 bytes per lane at scr2 to scr8, whatever the nominal 32 KiB, 128 KiB or 1 MiB. The named chip (cache mirror plus recompute) keeps that in SRAM at under 5 percent of its mirror and its gain does not move at any share under the 6 GB cap. | NOT sound as an anti-chip layer |
|
||||
| 3 | Is the verifier's one-warp simulation exact? | Exact when the GPU's lazy per-unit tag is unique over the arena's life and the arena holds no stale tag. The kernels rely on this and neither host guarantees it (no clear at allocation, no clear at the 32-bit wrap of the tag counter, 16.4 minutes on a 5090). With the host contract of section 4.3 the simulation is exact: 14 edge packs twice, 200 fuzz packs, consecutive units on one warp and the wrap inside a launch all match the CPU on Metal (228 of 228); a broken tag and a broken fill are caught (3 of 3). | sound with a host contract; today it is luck |
|
||||
| 4 | The attack surface of the writes | Out of bounds: impossible by the mask, 42 of 42 emitted kernels pass the static check, which catches six deliberate breaks. Aliasing: none, lane-major arenas disjoint by (warp, lane), two logical units of a wave64 get two arenas. Ordering: one lane, one slot, program order; no cross-lane sharing, no atomics needed. Alignment: 16-byte slots at 16-byte offsets from a 256-byte-aligned base. Wrap: identical to the CPU, tested at the launch level. | sound |
|
||||
| 5 | What a conformance vector must carry | The class and geometry, the fill and rewrite, the host contract (tags, clearing, groups a multiple of warps), two consecutive units on one warp with a forced slot collision, a unit in the top 256 nonces with the wrap inside the launch, and the fingerprint declared independent of the warp count. The standard three-unit vectors catch a broken tag only through base 1,000,000 and would miss it at a 1 MiB scratch. | defined in section 6 |
|
||||
|
||||
Recommendation (section 10): do not adopt layer 3 as the plan states it (read-modify-writes taken from the 16
|
||||
dataset loads). It replaces latency-bound dataset reads with cache-bound ones for the GPU, costs the named chip
|
||||
nothing it cannot keep in a few megabytes of SRAM, and leaves that chip's gain at 2.4x at every share. The lever
|
||||
that moves that chip is the mixer multiplier of the M16 analysis (x2 brings it to 1.2x, x4 to 0.6x, under the
|
||||
verifier's 10 ms gate). If a scratch is kept for another reason, add the read-modify-writes beside the 128 loads,
|
||||
never in their place, and ship the host contract and the vector of section 6 with it.
|
||||
|
||||
## 1. What the branch implements
|
||||
|
||||
| Piece | Where | What |
|
||||
|---|---|---|
|
||||
| Class | `igneum-pow/src/generator.rs:170-230` | `LoadClass { scratch: Some(k), scratch_kb }`: `k` of the 16 memory slots are `Op::Scratch`; `scratch_kb` KiB per warp of 16-byte slots, lane-major, `slots = kb x 2` per lane (32 KiB: 64, 128 KiB: 256); `scratch_slot_mask() = slots - 1` |
|
||||
| Draw | `generator.rs:488-491` | the first `k` of the 16 drawn load slots become scratch ops (a uniform k-subset); the source register follows the fresh-source rule like a load |
|
||||
| Fill | `igneum-pow/src/verify.rs:30` | `scratch_fill(seed, base, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot x 0x9e3779b1 + (j + 1) x 0x85ebca77)`, j in 0..2 |
|
||||
| Fold | `verify.rs:18` | `x = dst ^ w0; x = (rotl(x, 11) x 0x9e3779b1) ^ w1; x = (rotl(x, 11) x 0x9e3779b1) ^ w2; dst = x` (the read-width fold over the three data words) |
|
||||
| Rewrite | `verify.rs:41` | the slot becomes `(x ^ w1, rotl(x, 7) ^ w2, x + w0)` |
|
||||
| CPU model | `verify.rs:48-100`, `:305-312` | `ScratchModel`: per (lane, slot) a written bit and three words; an unwritten slot reads as its fill; one model per unit, so a unit starts from the fill |
|
||||
| Acceptance | `igneum-pow/src/accept.rs:202-215, 374` | a scratch site that reads one slot in all 32 lanes rejects the program (lane-constant site); scratch slots carry bit 31 in the address list and are left out of the distinct-address bound, which now covers the dataset loads only |
|
||||
| GPU statement | `igneum-pow/src/emit.rs:143-157` | `s_ = rN & mask; v_ = 16-byte load of slot s_; m_ = (v_.x == tag) ? ~0 : 0; w = (v_.yzw & m_) \| (fill & ~m_); fold; dst = x_; 16-byte store of (tag, x_ ^ w1_, rotl(x_, 7) ^ w2_, x_ + w0_)` in Metal, CUDA and OpenCL |
|
||||
| Persistent prologue | `emit.rs:159-175` | `lane = tid & 31; warp_ = tid >> 5; arena = scratch + (warp_ x 32 + lane) x words_per_lane; for (g_ = warp_; g_ < groups; g_ += nwarps_) { gbase = baseNonce + g_ x 32; tag = salt + g_; ... }` |
|
||||
| Hosts | `proto-metal/packbench.swift:144-164`, `proto-opencl/host.c:1025-1033, 1268` | the arena is allocated and never written by the host; `salt` starts at 1 and advances by the launch's unit count; no clear at allocation, none at the wrap |
|
||||
|
||||
The constraint of the night (coordinator, 5 October 2026): the whole working set on an 8 GB card stays under 6 GB
|
||||
(1 GiB table, the layer 5 hot table, the scratch of every resident warp, buffers), which caps the scratch at tens of
|
||||
KiB per warp. On an RTX 5090 at full occupancy (170 SMs x 64 warps = 10,880 warps, approximate hardware maximum;
|
||||
the measured version 2 kernel ran 24 warps per SM, 4,080 warps, `docs/bench-log.md` M11, 4 October 2026):
|
||||
|
||||
| Scratch per warp | 10,880 warps | 4,080 warps (measured occupancy) | Table + scratch at 10,880 | Under 6 GB with a 1 GiB table |
|
||||
|---|---|---|---|---|
|
||||
| 32 KiB | 340 MiB | 128 MiB | 1,364 MiB | yes |
|
||||
| 128 KiB | 1,360 MiB | 510 MiB | 2,384 MiB | yes |
|
||||
| 1 MiB (the first experiment) | 10,880 MiB | 4,080 MiB | 11,904 MiB | no |
|
||||
|
||||
## 2. Question 1: uniformity of what is written
|
||||
|
||||
### 2.1 As functions
|
||||
|
||||
The fill of word j of slot s for lane nonce n is `splitmix32(((n ^ seed[j]) + s x 0x9e3779b1 + (j + 1) x 0x85ebca77))`.
|
||||
`splitmix32` is a bijection of its 32-bit input; for fixed (seed, s, j) the input is a bijection of n. So over any
|
||||
2^32 consecutive nonces every 32-bit value appears once as the fill of (s, j): uniform. Test
|
||||
`fill_is_a_bijection_of_the_nonce`: 2^16 consecutive nonces give 2^16 distinct words for 7 slots x 3 word
|
||||
positions; the fill of lane l at base b equals the fill of lane 0 at base b + l; it wraps with the nonce
|
||||
(base 0xffffffe0, lane 32 equals nonce 0).
|
||||
|
||||
The rewrite `(x ^ w1, rotl(x, 7) ^ w2, x + w0)` is, for fixed old content w, a bijection of the fold value x in
|
||||
EACH word. Test `rewrite_is_a_bijection_of_the_fold_value`: 2^16 consecutive x give 2^16 distinct words in each
|
||||
position for 16 random w. Consequence: a uniform x gives a uniform word in every position, and the three words
|
||||
are three images of the same x, so a rewritten slot carries exactly 32 bits of new state behind 96 bits of
|
||||
storage (from w and any one written word, x is recovered; the test checks all three inversions).
|
||||
|
||||
The fold value x is `fold(dst, w)`, a bijection of `dst` for fixed w (xor, then rotate-multiply-xor twice; the
|
||||
multiplier is odd). So the written words are uniform whenever `dst` is, and `dst` is a register of the running
|
||||
program.
|
||||
|
||||
### 2.2 The attack: what a chip can precompute
|
||||
|
||||
The fill is a pure function of (seed, nonce, slot): precomputable, and meant to be (the verifier computes it too).
|
||||
A chip never stores a fill; it computes it in about 10 integer operations when a slot is first touched. The written
|
||||
words depend on `dst`, the register state at that instruction, which depends on every earlier instruction of the
|
||||
hash, including the dataset loads. Nothing about them is precomputable before the hash runs. This is the whole of
|
||||
what question 1 can give: the writes are as unpredictable as the registers. What that is worth is question 2.
|
||||
|
||||
### 2.3 The stats run (the `TESTS.md` section 3 shape)
|
||||
|
||||
Test `written_words_unbiased_and_rehit_rates`, M5 Max, 5 October 2026, `cargo test --test scratch`: for each
|
||||
class, programs of `igneum-genesis`, `igneum-genesis/stats1`, `igneum-genesis/stats2`, 2^11 units each (196,608
|
||||
hashes per class), closed-form dataset, every read-modify-write traced (`verify::interpret_warp_scratch`). Ones
|
||||
count per bit of every written word and of the change each rewrite makes (written XOR read), sigma = sqrt(N)/2,
|
||||
limit 6 sigma like the acceptance rule's output check.
|
||||
|
||||
| Class | Slots per lane | RMW per hash per lane | Rewrites traced | Max bias, written words (sigma) | Max bias, written XOR read (sigma) |
|
||||
|---|---|---|---|---|---|
|
||||
| scr2k32 | 64 | 16 | 3,145,728 | 2.61 | 3.40 |
|
||||
| scr4k32 | 64 | 32 | 6,291,456 | 2.18 | 3.81 |
|
||||
| scr8k32 | 64 | 64 | 12,582,912 | 3.63 | 2.25 |
|
||||
| scr2k128 | 256 | 16 | 3,145,728 | 3.36 | 2.19 |
|
||||
| scr4k128 | 256 | 32 | 6,291,456 | 2.71 | 3.68 |
|
||||
| scr8k128 | 256 | 64 | 12,582,912 | 2.73 | 2.60 |
|
||||
|
||||
576 bit positions (6 classes x 3 words x 32 bits) at under 4 sigma is what fair coins give. Verdict: no structural
|
||||
bias in what is written. Like `TESTS.md` section 3 this is a sanity check, not a proof of strength.
|
||||
|
||||
## 3. Question 2: no short cut avoids the writes
|
||||
|
||||
### 3.1 Inside a unit: the chain is dependent
|
||||
|
||||
Slot s of lane l, touched d times in a unit, holds `w_d = rewrite(x_d, w_{d-1})`, `w_0 = fill`, with
|
||||
`x_i = fold(dst_i, w_{i-1})`. `x_i` depends on the slot content before it, which depends on every earlier fold
|
||||
value of that slot; and `dst_i` is the register state, which the earlier fold values entered. Test
|
||||
`slot_is_replayable_from_its_fold_values`: a slot after 64 read-modify-writes is reproduced from the fill and the
|
||||
64 fold values; dropping one diverges. So a chip cannot skip a write and still read the slot later. It has three
|
||||
ways to hold a slot, all exact:
|
||||
|
||||
| Store | Bytes per lane | Cost on a re-hit |
|
||||
|---|---|---|
|
||||
| Dense: every slot, 12 data bytes plus a valid bit | 12 x slots: 776 (64 slots), 3,104 (256), 24,832 (2,048) | one SRAM read |
|
||||
| Sparse: only touched slots, 12 bytes plus a slot index | about 13 x distinct: 185 to 820 (table below) | one lookup |
|
||||
| Implicit: only the fold values, 4 bytes plus a slot index per read-modify-write, replay on a re-hit | 5 x 8k: 80 (scr2), 160 (scr4), 320 (scr8) | d rewrites of 5 integer ops |
|
||||
|
||||
The implicit store is smaller than the dense one whenever `slots > 8k / 3`: at scr4 above 10.7 slots, at scr8
|
||||
above 21.3. So "the smallest scratch at which keeping it implicitly is dearer than storing it" is 8k/3 slots per
|
||||
lane, 2.7 to 5.3 KiB per warp at scr4 to scr8. Every size on the table, 32 KiB and above, is past it: a chip
|
||||
keeps the scratch implicitly in 80 to 320 bytes per lane at any nominal size, and the replay cost is bounded by
|
||||
the re-hit depth, which the next table measures.
|
||||
|
||||
### 3.2 The re-hit rate at 64 and 256 slots (and at 2,048)
|
||||
|
||||
Measured in the same test run (every read-modify-write of 196,608 hashes per class traced; a re-hit is a read of a
|
||||
slot the same unit wrote earlier). Birthday: `distinct = S (1 - (1 - 1/S)^n)` for n uniform draws from S slots.
|
||||
|
||||
| Class | S | n = RMW per hash | Distinct slots, birthday | Re-hits, birthday | Re-hit %, birthday | Re-hit %, measured | Max chain depth seen | Slot histogram against uniform |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| scr2k32 | 64 | 16 | 14.26 | 1.74 | 10.9 | 12.58 | 7 | chi2 z 22,023; hottest slot 2.74x, coldest 0.83x |
|
||||
| scr4k32 | 64 | 32 | 25.33 | 6.67 | 20.8 | 21.47 | 8 | z 10,880; 1.87x, 0.92x |
|
||||
| scr8k32 | 64 | 64 | 40.64 | 23.36 | 36.5 | 36.99 | 9 | z 7,587; 1.39x, 0.91x |
|
||||
| scr2k128 | 256 | 16 | 15.54 | 0.46 | 2.9 | 3.84 | 5 | z 19,146; 5.10x, 0.82x |
|
||||
| scr4k128 | 256 | 32 | 30.14 | 1.86 | 5.8 | 6.25 | 6 | z 9,632; 3.06x, 0.90x |
|
||||
| scr8k128 | 256 | 64 | 56.72 | 7.28 | 11.4 | 11.89 | 6 | z 5,873; 1.98x, 0.90x |
|
||||
| 1 MiB (not run) | 2,048 | 32 | 31.76 | 0.24 | 0.8 | | | |
|
||||
|
||||
Two readings. First, the slot a read-modify-write addresses is the low 6 or 8 bits of a program register, and
|
||||
those bits are not uniform: `or` sets them, `mul` clears them, so one slot of 256 is addressed 5.1 times as often
|
||||
as the mean and the re-hit rate runs 2 to 33 percent above the birthday rate. For the dataset the same bias on the
|
||||
low bits of a 28-bit address is harmless (it moves the read inside an item); for a 64-slot scratch it concentrates
|
||||
the chain. Second, the chain depth is small: at scr4k32 the deepest slot in 196,608 hashes saw 8 earlier
|
||||
read-modify-writes; a replay costs at most 8 x 5 integer operations, against about 1,170 for one dataset item.
|
||||
|
||||
### 3.3 The live state is bounded by the read-modify-write count, not by the size
|
||||
|
||||
The verifier evaluates one unit from nothing but (program, day, nonce group): `ScratchModel::new` per unit,
|
||||
`verify.rs:296`. Every conforming GPU must therefore start every unit from the fill, which the tag does
|
||||
(section 4). So no state crosses a unit boundary, and the state a unit can ever read back is what it wrote itself:
|
||||
at most 8k slots per lane. The nominal size only sets how often those 8k writes land on the same slot (the table
|
||||
above). The scratch's "memory" is 8k x 16 bytes per lane of touched slots, 256 bytes to 1 KiB at scr2 to scr8,
|
||||
and a chip holds it implicitly in 80 to 320 bytes.
|
||||
|
||||
The attack of rolling back or sharing scratch between units has nothing to take: a unit starts from the fill
|
||||
whatever ran before it, so a chip that clears 64 valid bits per unit has rolled back, and nothing one unit wrote
|
||||
is readable by another. The CPU verifier is that chip.
|
||||
|
||||
### 3.4 The named chip, and what the scratch costs it
|
||||
|
||||
The strongest chip the plan has priced (coordinator, 5 October 2026): the whole 256 MiB cache on the die, computing
|
||||
every dataset item on the fly. Its cache SRAM, from `docs/analysis/sram-mirror.md` revision 2 (`ca2-analysis`
|
||||
e6085c6), headline at shipped-product density / bit-cell lower bound, dollars per good die approximate: 164 / 83
|
||||
mm^2 and $30 / $13 at N7 (shipped density from AMD 3D V-Cache, 64 MB on 41 mm^2, Hot Chips 2021); 128 / 64 mm^2 and
|
||||
$46 / $21 at N5, N3E and Intel 18A (TSMC N5 HD macro 31.8 Mib/mm^2 after assist overhead, SemiAnalysis, December
|
||||
2022); 106 / 54 mm^2 and $56 / $26 at N2; with a 96 MB hot table 226 / 114 at N7, 175 / 89 at N5, 146 / 74 at N2.
|
||||
The chip's cache cost in the table below is the N5 headline, 128 mm^2 and $46 per good die. It computes every item
|
||||
through the mixer (`docs/analysis/m16-recompute-attacker-2026-10-05.md`: 128 items per hash, about 1,170 integer
|
||||
operations per item, 150,000 per hash; at a 50 T op/s integer budget equal to a 5090's, approximate, 0.33 Ghash/s).
|
||||
Against the measured version 2 rate of the RTX 5090, 139.7 MH/s (`docs/bench-log.md` M11, 4 October 2026), that is
|
||||
2.4x before any fixed-function factor, 7x with the 3x the M16 analysis allows (approximate).
|
||||
|
||||
Units in flight on that chip. It has no DRAM latency to cover: every one of its 1,024 cache reads per hash is an
|
||||
on-die SRAM read. Its hash latency is the dependent chain: 128 items x (8 dependent SRAM reads plus 9 mixer
|
||||
applications). At about 10 ns per on-die read and about 40 ns per 130-operation mixer on a 16-wide integer
|
||||
pipeline at 2 GHz (both approximate), an item is about 0.4 us and a hash about 50 us; at 0.33 Ghash/s that is
|
||||
about 17,000 hashes in flight, 530 units of 32 lanes. A tighter pipeline halves it. The GPU covers DRAM latency (40 to 48 ns row
|
||||
cycle, MEMSYS 2018, more under load) with 130,560 lanes in flight at the measured occupancy (4,080 warps x 32), 348,160 at
|
||||
full occupancy, that is 8 to 20 times more lanes than the chip needs.
|
||||
|
||||
What the scratch costs that chip, per variant, with the arithmetic:
|
||||
|
||||
Chip cache mirror: 128 mm^2, $46 per good die (N5 headline; 64 mm^2, $21 bit-cell lower bound). Chip scratch SRAM at
|
||||
the same two densities (2.1 MB/mm^2 headline, 4.2 MB/mm^2 lower bound at N5):
|
||||
|
||||
| Variant | Dataset loads per hash | Chip ops per hash | Chip rate at 50 T op/s | 5090 rate | Chip gain | Chip scratch SRAM at 17,000 lanes, implicit store | Same, dense 64-slot store | Dense store as mm^2, headline / lower bound (N5) | Share of the 256 MiB mirror (any density) |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| scr0 (control), 128 loads | 128 | 150,000 | 333 MH/s | 139.7 measured | 2.4x | 0 | 0 | 0 | 0 |
|
||||
| 12.5% replaced (scr2) | 112 | 131,400 | 381 | 160 projected (128/112 x 139.7) | 2.4x | 1.4 MB | 13 MB | 6.2 / 3.1 mm^2 | 4.9% |
|
||||
| 25% replaced (scr4) | 96 | 112,800 | 443 | 186 projected | 2.4x | 2.7 MB | 13 MB | 6.2 / 3.1 | 4.9% |
|
||||
| 50% replaced (scr8) | 64 | 75,600 | 661 | 279 projected | 2.4x | 5.4 MB | 13 MB | 6.2 / 3.1 | 4.9% |
|
||||
| 12.5% added (16 RMW beside 128 loads) | 128 | 150,200 | 333 | 139.7 or below | 2.4x or more | 1.4 MB | 13 MB | 6.2 / 3.1 | 4.9% |
|
||||
| 25% added | 128 | 150,400 | 332 | 139.7 or below | 2.4x or more | 2.7 MB | 13 MB | 6.2 / 3.1 | 4.9% |
|
||||
| 50% added | 128 | 150,800 | 332 | 139.7 or below | 2.4x or more | 5.4 MB | 13 MB | 6.2 / 3.1 | 4.9% |
|
||||
| 256-slot dense store (128 KiB class), any share | | | | | | | 53 MB | 25 / 12.6 | 20% |
|
||||
|
||||
How the rows are computed: a read-modify-write costs the chip about 12 integer operations (fold and rewrite) and
|
||||
one SRAM access; replacing a load removes an item derivation (1,170 operations); the 5090's rate for a replaced
|
||||
load is projected from the measured distinct-load bound (the card's rate tracks distinct dataset loads per hash,
|
||||
`docs/bench-log.md` 3 October, 23.7 G loads/s at 1 GiB; the readwidth agent's M5 Max measurement of the night,
|
||||
relayed by the coordinator, shows the same: 27.7 MH/s at v2 to 29.4-31.7 at 25 percent replaced and 44.4-49.1 at 50
|
||||
percent, 32 KiB per warp). The scratch SRAM is 17,000 lanes x 80 to 320 bytes (implicit) or x 776 bytes (dense at
|
||||
64 slots) or x 3,104 bytes (dense at 256 slots); its share of the mirror is a ratio of bytes, 4.9 or 20 percent,
|
||||
whichever density is used for both; the implicit store (the chip's cheaper choice at every size, section 3.1) is
|
||||
0.5 to 2 percent. The chip's gain is set by operations per dataset item and the GPU's distinct-load bound, and the
|
||||
scratch touches neither.
|
||||
|
||||
Plain answer to the coordinator's question: no read-modify-write share under the 6 GB cap, replaced or added,
|
||||
brings the named chip under 2x. The share would be chosen as the smallest at which the chip falls under 1.5x, and
|
||||
there is none: the gain is 2.4x at 0, 12.5, 25 and 50 percent, 32 or 128 KiB. This changes nothing about the public
|
||||
claim that layer 3 would have changed: the claim must rest on the mixer, not on the scratch.
|
||||
|
||||
The lever that does move that chip, from the M16 table, beside it:
|
||||
|
||||
| Mixer cost multiplier | Chip ops per hash | Chip rate | Gain against 139.7 MH/s, no fixed-function factor | With a 3x factor (approximate) | CPU verify per warp (M16 table, scaled from 0.41 to 1.2 ms) | 5090 daily dataset build |
|
||||
|---|---|---|---|---|---|---|
|
||||
| x1 (today) | 150,000 | 333 MH/s | 2.4x | 7.2x | 0.4 to 1.2 ms | 13.4 ms |
|
||||
| x2 | 300,000 | 167 | 1.2x | 3.6x | 0.8 to 2.4 ms | 27 ms |
|
||||
| x4 | 600,000 | 83 | 0.6x | 1.8x | 1.6 to 4.8 ms | 54 ms |
|
||||
| x8 | 1,200,000 | 42 | 0.3x | 0.9x | 3.3 to 9.6 ms | 107 ms |
|
||||
|
||||
The mixer multiplier leaves the honest hash rate untouched (the miner pays the mixer once a day), costs the chip
|
||||
linearly, and is bounded by the 10 ms verification gate (x8 is at the gate's edge on this core, and the 2019-class
|
||||
core of O-1.14 is unmeasured). The scratch costs the honest GPU a measured share of its rate when it spills the
|
||||
cache and nothing when it does not, and costs the chip a few megabytes. The comparison is not close.
|
||||
|
||||
### 3.5 Where the GPU's writes would cost DRAM latency, and why that does not help
|
||||
|
||||
The GPU's hot scratch footprint is not the nominal size either: it is the slots in-flight units have touched,
|
||||
about `warps x 32 lanes x distinct slots x 16 bytes` (x 2 at a 32-byte sector, approximate): on the 5090 at 4,080
|
||||
resident warps and scr4, 25.3 slots at 64 or 30.1 at 256, 53 to 63 MB of slots, 100 to 125 MB in sectors, around
|
||||
the card's 96 MiB L2 (`docs/bench-log.md`, 3 October). The readwidth agent's M5 Max rows (coordinator's message:
|
||||
the rate rises with the share at 32 and 128 KiB) show the scratch sitting in that chip's caches at 4,096 warps.
|
||||
To push the writes to DRAM latency the hot footprint must pass the last-level cache at the resident count:
|
||||
`96 MiB / 4,080 warps = 24 KiB per warp`, which at 512 bytes of touched slots per lane per read-modify-write slot
|
||||
means `8k x 512 B > 24 KiB`, k above 6 (above 48 read-modify-writes per hash) at ANY nominal size on the table, or
|
||||
a higher resident count. That fits the 6 GB cap (it is the hot set, not the arena, that matters), and it costs
|
||||
the honest miner a DRAM-latency read-modify-write per slot (a DRAM row cycle is 40 to 48 ns across DDR4, GDDR5 and
|
||||
HBM2, Li, Reddy and Jacob, MEMSYS 2018; the loaded latency a GPU kernel sees is higher, approximate; DRAM latency
|
||||
improved 1.3x in two decades while bandwidth improved 20x, Chang 2017, so no memory technology an attacker could
|
||||
buy removes it, and no shipped mining chip has used HBM or stacked memory) while the named chip still keeps the same
|
||||
hot set in a few megabytes of SRAM at 8 to 20 times fewer lanes in flight. The write path cannot be made to cost the
|
||||
chip more than the GPU, because the GPU must keep 8 to 20 times more of it live.
|
||||
|
||||
## 4. Question 3: the verifier's one-warp simulation is exact
|
||||
|
||||
### 4.1 Lazy fill on both sides
|
||||
|
||||
The CPU initialises lazily with a written bit per (lane, slot), one model per unit. The GPU initialises lazily with
|
||||
a 32-bit tag in word 0 of each 16-byte slot: a slot whose tag equals the unit's tag reads as written, any other
|
||||
reads as the fill (`emit.rs:143-157`). There is no explicit fill and no reset between units of a persistent warp
|
||||
(`emit.rs:159-175`: the loop over `g_` keeps the arena). The two agree if and only if, when a unit first touches a
|
||||
slot, that slot does not already carry the unit's tag. That is:
|
||||
|
||||
1. Tags are unique over the life of the arena's contents (`tag = salt + g_`, `salt` the host's running counter).
|
||||
2. The arena holds no word equal to a live tag in a slot's tag position before the unit writes it.
|
||||
|
||||
### 4.2 The attacks (the bug classes)
|
||||
|
||||
| Case | What happens | Today |
|
||||
|---|---|---|
|
||||
| Recycled allocation | A fresh process starts `salt` at 1 (`packbench.swift:144`, `host.c:1027`). If the driver hands back the previous process's arena with its contents (Metal, CUDA and OpenCL do not promise zeroed memory, approximate), slots tagged 1..N from the old run match the new run's first units exactly, and those units read stale words instead of the fill: a CPU mismatch on every colliding slot. | not guarded; passes on this Mac because fresh allocations read as zero in practice and tag 0 is never issued (luck, not contract) |
|
||||
| Tag counter wrap | `salt` is 32 bits and advances by units per launch. A 5090 at 139.7 MH/s runs 4.37 M units/s, 2^32 units in 984 s: the counter wraps every 16.4 minutes on one card (81.8 minutes on the M5 Max at 28 MH/s). After the wrap a slot whose LAST writer carried the repeated tag reads as written. With 10,880 arenas each slot is rewritten about 395,000 times between two uses of one tag (at 64 slots a unit leaves a slot untouched with probability 0.60; 0.60^395,000 is 0), so on a full card the wrap is harmless in practice; on a one-warp launch repeated 2^32 times it is not. | not guarded |
|
||||
| Tag 0 on zeroed memory | A host that starts `salt` at 0 gives unit 0 the tag 0, which a zeroed arena carries in every slot: unit 0 reads zeros for every first touch. | both hosts start at 1; nothing in the pack says they must |
|
||||
| `groups` not a multiple of the warp count | Warps run different trip counts; the OpenCL local-memory exchange path carries a barrier inside the loop (spec 1.9), so a short warp hangs or desynchronises. | `packbench` refuses it; `host.c` rounds the batch |
|
||||
|
||||
### 4.3 The host contract that makes the simulation exact
|
||||
|
||||
A host of a scratch class MUST: allocate the arena as `warps x 32 x words_per_lane` words and zero it; issue tags
|
||||
from a 32-bit counter that starts at 1 and advances by the unit count of every launch; zero the arena again before
|
||||
any launch whose tags would pass 2^32 - 1 (tag 0 is never issued); launch `groups` as a multiple of the warp count.
|
||||
The zeroing costs one memset of the arena (340 MiB at 32 KiB x 10,880 warps) every 2^32 units, 16 minutes on a
|
||||
5090. This is the class fix for all four rows: with it the GPU's tag test and the CPU's written bit are the same
|
||||
predicate.
|
||||
|
||||
### 4.4 The tests (Metal, M5 Max, 5 October 2026)
|
||||
|
||||
Two consecutive units on one persistent warp and the wrap inside a launch (`packbench --warps 1`,
|
||||
`--batch-base 4294967040`, the option added on this branch); the hand-built edge programs that force every
|
||||
read-modify-write of a hash onto one slot (so two consecutive units on one arena collide on every slot); the
|
||||
deliberate breaks. Results in section 7.2. On the CPU, the same edge programs against an independent hand model
|
||||
(a second interpreter with its own slot store, `tests/scratch.rs`): 56 of 56 cases match, and the hand model with
|
||||
its rewrite words swapped mismatches on every case (the comparison has teeth).
|
||||
|
||||
## 5. Question 4: the attack surface of the writes
|
||||
|
||||
| Surface | Argument | Test |
|
||||
|---|---|---|
|
||||
| Out of bounds | `s_ = rN & (slots - 1)`, so `s_ < slots`; the lane's arena is `(warp_ x 32 + lane) x 4 x slots` words from the base, the access is `arena + 4 x s_ + 0..3`, the largest index is `warps x 32 x 4 x slots - 1`, the host's allocation. The emitter has one scratch template (`emit.rs:143`) and it masks. | `scr_packs_regenerate_and_pass_the_static_scratch_check`: 42 of 42 emitted kernels (7 scr packs x 6 files, the OpenCL bound file carrying two kernels) regenerate byte for byte from program.json and pass the text check: k masked slot definitions with the class mask, k tagged stores, 3k fill calls, one arena definition with the class stride, one tag definition, no `scratch[`; six deliberate breaks caught (section 8) |
|
||||
| Aliasing between lanes | Lane-major: lane l of warp w owns words `[(32w + l) x 4S, (32w + l + 1) x 4S)`; two (w, l) pairs give disjoint ranges. Inside the range a slot is 4 words at `4 x s_`, so two slots of one lane are disjoint too. | the `lanevar` edge program: one init-dependent slot per lane, 32 lanes at 64 slots share slots in pairs by the birthday bound; any cross-lane aliasing would change the fold; 128 of 128 lanes on Metal (section 7.2) |
|
||||
| Wave64 (two logical units in one hardware wave) | `warp_ = tid >> 5`, so the two halves get `warp_ = 2w` and `2w + 1`, two arenas; `gbase` and `tag` are per `g_`, per half. | not run on wave64 hardware (the OpenCL emulator's persistent launch is on the readwidth commit; unverified here) |
|
||||
| Determinism: alignment | A slot is 16 bytes at byte offset `16 x (lane_base + s_)`; the arena base is the buffer base: Metal, CUDA and OpenCL allocations are at least 128-byte aligned (CUDA 256, OpenCL `CL_DEVICE_MEM_BASE_ADDR_ALIGN` at least the largest built-in type, approximate from memory), so every 16-byte vector access is aligned. | Metal: every run of section 7 |
|
||||
| Determinism: ordering | A lane's two read-modify-writes of the same slot in one hash are a load and a store, then a load and a store, from one thread to one address: program order within a thread holds in every model. No other thread touches the slot (aliasing row), so no atomics, fences or barriers are needed and none are emitted. | `slot0` and `sixteen` edge programs: 64 and 128 dependent read-modify-writes on one slot per lane per hash, standalone and as the second unit on a warp |
|
||||
| Determinism: vendors | The statement is integer only: xor, rotate by immediate, multiply, add, a 16-byte load and store. Bit-exact across Metal, CUDA and OpenCL by construction; measured only on Metal here. | Metal; CUDA and OpenCL runs are PC jobs (not mine tonight) |
|
||||
| 32-bit nonce wrap | `gbase = baseNonce + g_ x 32` and `nonce = baseNonce + gid` wrap in 32-bit arithmetic; `scr_fill(gbase + lane)` wraps like the CPU's `base.wrapping_add(lane)`; `out[gid]` indexes by launch position, not by nonce. An aligned unit never straddles 2^32 (spec 1.9), so the wrap case is a launch whose unit SEQUENCE crosses it. | `packbench --batch-base 4294967040 --batch-log2 9`: 16 units from 0xffffff00, the ninth at gbase 0; fingerprint identical at 1 and 4 warps (section 7.2); every fuzz pack runs that launch |
|
||||
|
||||
## 6. Question 5: what a vector for the scratch class must carry
|
||||
|
||||
Before a scratch pack can be a conformance vector (plan step 4, "only then a vector"), it must carry, beyond what
|
||||
`igneum-program-pack-3` carries today:
|
||||
|
||||
1. The class in the program id and the pack (`scr<k>k<kb>`: it is, `program_id_class`, `generator.rs:400-412`)
|
||||
and the geometry (slots per lane, words per lane, bytes per warp: it is, `program.h`).
|
||||
2. The fill and the rewrite as text (it is, `program.json` "scratch").
|
||||
3. The host contract of section 4.3 as text in `program.h` and `program.json`: tag counter from 1, zero at
|
||||
allocation and at the wrap, `groups` a multiple of the warp count. Not there today.
|
||||
4. Vectors that exercise the tag path, which the three standard units do not reliably: two consecutive units on
|
||||
one warp (bases 0 and 32 in one one-warp launch) for a program whose consecutive units collide on a slot. At
|
||||
64 slots any generated program collides (25 touched of 64 per unit; the broken-tag run of section 8 was caught by
|
||||
base 1,000,000, a warp's 16th unit, and NOT by a two-unit launch whose vectors lack base 32). At 2,048 slots two
|
||||
consecutive units share a touched slot with probability about 0.4 (32 x 32 / 2,048 expected overlaps = 0.5), so
|
||||
the standard vectors would miss a broken tag at the 1 MiB size with probability about 0.6 per unit pair. The
|
||||
edge programs `slot0` and `sixteen` collide on every slot at every size: a vector set should carry one.
|
||||
5. A unit in the top 256 nonces with the launch crossing 2^32 (`--batch-base` near the top, at least two warps).
|
||||
6. The batch fingerprint declared independent of the warp count (`8c07620f4d9adefd` for scr4k32 at 2^12 nonces
|
||||
from base 0 at 1, 2 and 128 warps, section 7.2): unit independence is the property the per-unit reset gives, and
|
||||
a fingerprint that moved with the warp count would mean a unit read another unit's slot.
|
||||
|
||||
## 7. Tests and results
|
||||
|
||||
### 7.1 CPU (`igneum-pow/tests/scratch.rs`, `cargo test -j4 --test scratch`, M5 Max, 5 October 2026, 3.6 s)
|
||||
|
||||
| Test | What | Result |
|
||||
|---|---|---|
|
||||
| `rewrite_is_a_bijection_of_the_fold_value` | 16 random slot contents x 2^16 consecutive fold values, each written word distinct; the three inversions | pass |
|
||||
| `fill_is_a_bijection_of_the_nonce` | 7 slots x 3 words x 2^16 nonces distinct; lane and base interchange; wrap | pass |
|
||||
| `written_words_unbiased_and_rehit_rates` | 6 classes x 3 seeds x 2^11 units, every rewrite traced: bias within 6 sigma (worst 3.63), re-hit rate within 0.9x to 2x of birthday, slot histogram, depth histogram | pass (tables of sections 2.3 and 3.2) |
|
||||
| `edge_programs_match_the_hand_model` | 7 edge programs x 2 geometries x 4 bases (0, 32, 0x7ffffff0, 0xffffffe0) against an independent hand model; the slots driven and the re-hit counts as built; the mutated hand model mismatches | 56 of 56 pass, 56 of 56 teeth |
|
||||
| `scr_packs_regenerate_and_pass_the_static_scratch_check` | 7 scr packs: program and program id from program.json, 6 kernel texts byte for byte, static scratch check on all 42, the pack's vectors from the CPU; six deliberate breaks caught | pass |
|
||||
| `fuzz_scr_programs_cpu` | 200 generated programs over the six classes, generator contract and acceptance on every one, 4 units each (one in 0..224, one around 2^31, one in the top 256 nonces, one uniform), traced run equal to the untraced run, every slot inside the lane; writes the 214 packs for Metal with `IGNEUM_SCRATCH_PACKS_OUT` | pass; 200 of 200 have a unit in the top 256 |
|
||||
| `slot_is_replayable_from_its_fold_values` | 64 dependent read-modify-writes replayed from the fill and the fold values; one dropped diverges | pass |
|
||||
|
||||
The rest of the crate: 33 of 34 lib tests and all pack tests pass; `verify::tests::fold_and_wide_fetch` fails on the
|
||||
readwidth tip itself (`verify.rs:508`, `k as u32 * 0x9E37_79B1` overflows under the test profile's overflow
|
||||
checks; the readwidth agent's test, reported to its owner, not touched here).
|
||||
|
||||
### 7.2 Metal (`proto-metal/packbench` built from this branch, M5 Max, 5 October 2026, under `with-lock.sh run`)
|
||||
|
||||
| Run | Launch | Expected | Result |
|
||||
|---|---|---|---|
|
||||
| scr4k32, standard pack | 2,048 warps, 2^24 nonces, 1 batch | 3 of 3 standalone, 3 of 3 in batch | PASS, fingerprint `3d1af881bd978fb9`; 1.8 s wall for the whole run (compile, cache, 1 GiB build, vectors, batch) |
|
||||
| scr4k32, warp-count independence | 2^12 nonces (128 units) at 1, 2 and 128 warps | one fingerprint | `8c07620f4d9adefd` at all three, PASS |
|
||||
| scr4k32, wrap inside the launch | 512 nonces from 0xffffff00 at 1 and 4 warps | one fingerprint, the base-0 vector inside the window after the wrap | `8e9e233234d3a297` at both, in-batch 1 of 1, PASS |
|
||||
| scr4k32, broken tag (`tag = salt`), standard vectors | 2,048 warps, 2^24 | the base-1,000,000 vector (warp 530's 16th unit) fails | standalone 3 of 3, in batch 2 of 3, overall FAIL (caught) |
|
||||
| scr4k32, broken tag, two units on one warp | 1 warp, 2^6 | nothing to catch it: the standard vectors have no base 32 | standalone 3 of 3, in batch 1 of 1, PASS (missed: the point of section 6 item 4) |
|
||||
| 14 edge packs (7 programs x 32 and 128 KiB), run A | 1 warp, 2^6 (units at bases 0 and 32 on one arena) | 4 of 4 standalone, 2 of 2 in batch each | 14 of 14 PASS (56 of 56 standalone units, 28 of 28 in batch) |
|
||||
| 14 edge packs, run B | 1 warp, 2^9 from 0xffffff00 (16 units on one arena, the wrap inside) | 4 of 4 standalone, 3 of 3 in batch each | 14 of 14 PASS (56 of 56, 42 of 42) |
|
||||
| edge `slot0` at 32 and 128 KiB, broken tag (`tag = salt`) | 1 warp, 2^6 | standalone 4 of 4, in batch 1 of 2, FAIL | as expected at both geometries: the second unit read the first's slot 0 and FAILED; the standalone units passed |
|
||||
| edge `slot0` at 32 KiB, broken lazy fill (`m_` forced to all ones: a first touch reads the stale words) | 1 warp, 2^6 | standalone fails | 0 of 4 standalone, 0 of 2 in batch, FAIL (lane 0 of base 0: GPU `64b49aeb987dae69`, expected `9ff3a2021f66b5be`) |
|
||||
| 200 fuzz packs (scr2k32 29, scr4k32 26, scr8k32 42, scr2k128 32, scr4k128 26, scr8k128 45; datasets 64 MiB, 256 MiB, 1 GiB) | 2 warps, 2^9 from 0xffffff00 (8 units per warp, the wrap inside) | 4 of 4 standalone, 2 of 2 in the window, 200 of 200 PASS | 200 of 200 PASS: 800 of 800 standalone units (25,600 hashes), 400 of 400 in batch; 91 s for the 200 runs |
|
||||
|
||||
Totals on Metal: 228 of 228 runs PASS where a pass was expected, 3 of 3 FAIL where a failure was built in.
|
||||
|
||||
## 8. Deliberate breaks (the watcher rule)
|
||||
|
||||
| Break | Where | Caught by | Evidence |
|
||||
|---|---|---|---|
|
||||
| One slot mask dropped (Metal) | copy of scr4k32 `program.metal` | static check: "masked slot followed by the load: 3, expected 4" | test output |
|
||||
| Mask 63 changed to 127 on every RMW (Metal) | same | "masked slot followed by the load: 0, expected 4" | test output |
|
||||
| Arena stride 256 changed to 128 words (Metal) | same | "arena definition: 0, expected 1" | test output |
|
||||
| A stray `arena[0]` and `scratch[1]` access (Metal) | same | "arena mentions: 10, expected 9; direct scratch indexing: 1, expected 0" | test output |
|
||||
| One slot mask dropped (OpenCL, CUDA) | copies of scr4k32 `kernel.cl`, `kernel.cu` | "masked slot followed by the load: 3, expected 4" | test output |
|
||||
| Wrong class geometry or RMW count or kernel count passed against a right text | the same text | the check fails | test output |
|
||||
| `tag = salt` (every unit of a launch shares the tag) | copy of scr4k32 `program.metal`, on the GPU | the base-1,000,000 vector in a 2,048-warp batch | `vectors standalone 3/3, in batch 2/3`, overall FAIL |
|
||||
| the same on the `slot0` edge pack, two units on one warp | on the GPU | in-batch 1 of 2 | bench-log entry |
|
||||
| lazy fill broken (`m_` all ones) | copy of the `slot0` edge pack, on the GPU | standalone vectors | bench-log entry |
|
||||
| The hand model's rewrite words swapped | `tests/scratch.rs` | every edge case mismatches | 56 of 56 |
|
||||
|
||||
The out-of-bounds break (mask dropped) was not run on the GPU on purpose: Metal does not bounds-check device
|
||||
buffers (`TESTS.md` section 5), so a run would read another lane's or another buffer's words and "did not crash" would
|
||||
prove nothing. The static check is the guard, as it is for the dataset mask.
|
||||
|
||||
## 9. What is unverified
|
||||
|
||||
1. CUDA and OpenCL runs of the scratch packs on NVIDIA and AMD (PC jobs, reserved for the readwidth agent tonight);
|
||||
the 5090's rate per variant, so the "projected" column of section 3.4 is the distinct-load bound, not a
|
||||
measurement. Wave64 hardware for the two-arena argument.
|
||||
2. The chip-side latency figures of section 3.4 (10 ns SRAM read, 40 ns mixer) are approximate; the conclusion
|
||||
does not depend on them: at ten times the in-flight count the scratch is still under a sixth of the mirror.
|
||||
3. The recycled-allocation case was not reproduced (it needs a driver that hands back live contents); the argument
|
||||
is that nothing forbids it and the contract of 4.3 removes it.
|
||||
4. The slot-bias finding (section 3.2) was measured on three seeds per class; the hottest-slot ratio will vary by
|
||||
program.
|
||||
5. `verify::tests::fold_and_wide_fetch` on the readwidth tip (section 7.1).
|
||||
|
||||
## 10. Recommendation
|
||||
|
||||
1. Layer 3 is sound as a construct: the written words are uniform, the chain inside a unit has no short cut, the
|
||||
kernels cannot write out of bounds, and with the host contract of section 4.3 the CPU's one-warp simulation is
|
||||
exact (14 edge packs, 200 fuzz packs, the wrap, consecutive units on one arena, on Metal).
|
||||
2. Layer 3 is not sound as a chip-resistance layer, at the capped size or at any size: CPU verification resets the
|
||||
scratch per unit, so its live state is 8k slots per lane whatever the arena, a chip keeps it implicitly in 80 to
|
||||
320 bytes per lane, and the named chip (on-die cache mirror plus recompute) keeps its whole scratch in 1.4 to
|
||||
13 MB of SRAM at 530 units in flight, 3 to 5 percent of its mirror. Its gain stays at 2.4x (7x with a 3x
|
||||
fixed-function factor, approximate) at 0, 12.5, 25 and 50 percent, replaced or added, 32 or 128 KiB. No share
|
||||
under the 6 GB cap brings it under 2x.
|
||||
3. Taking the read-modify-writes from the 16 dataset loads makes the hash less memory-hard for everyone: the GPU
|
||||
measured faster at every share on the M5 Max (readwidth rows), and the chip's operations per hash fall with the
|
||||
loads. If a scratch is kept at all, add it beside the 128 loads. There is no reason found here to keep one.
|
||||
4. The lever that moves the named chip is the M16 mixer multiplier: x2 to 1.2x, x4 to 0.6x against the measured
|
||||
5090 rate, at 0.8 to 4.8 ms of verification per warp against the 10 ms gate. Decision 2 should price that
|
||||
against the gate on the 2019-class core (O-1.14) rather than layer 3.
|
||||
5. If Josh keeps layer 3 for a reason outside this analysis: ship the host contract in the pack, add the four
|
||||
vector items of section 6 (consecutive units with a forced collision, the wrap launch, the warp-count-independent
|
||||
fingerprint, the contract text), and run the CUDA and OpenCL twins of section 7.2 on the PCs before the class
|
||||
becomes a genesis rule.
|
||||
|
|
@ -1643,3 +1643,25 @@ Reading: on this GPU a signed-byte dot4 costs 4.7 ALU-chain steps and an unsigne
|
|||
| gfx1036 (integrated RDNA 2, 2 CUs), 3683.0 | 40.6 (105.9) | 15.8 (272.3), 2.6x | `sudot4` does not build: "needs target feature dot8-insts" | not listed | alu and dot4e yes |
|
||||
|
||||
Reading: one `dp4a` on the 5090 costs about one ALU-chain step (7.45 T dot4/s, 0.85 of the chain's 8.75 T steps/s); one `v_dot4_i32_iu8` on the 9070 XT the same (0.66 T, 0.95 of its chain). Emulating the signed dot4 costs 6.0x the instruction on NVIDIA (the OpenCL compiler does not fold the four sign-extended products into `dp4a`) and 1.38x on AMD. Vendor ratios: the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on hardware dot4; against the M5 Max's best (unsigned emulation, 0.55 T) it is 10x on the chain and 13.6x on dot4. The hash itself is bound by DRAM reads, so these per-op numbers bound a family's cost and are not hash rates (`docs/analysis/int8-matrix-family.md` section 4). Adrenalin's OpenCL C accepts the clang builtin and emits the instruction on RDNA 4 (the third-party RDNA 3 report of the same route is now confirmed on this card); no PC platform lists the Khronos integer-dot extension. The 5090 SM clock read 2,505 MHz before and after (nvidia-smi; 2,850 MHz while mining in the telemetry entry), so the card was idle for the probe.
|
||||
## 5 October 2026, layer 3 scratch soundness (Counter ASIC 2.0 step 4; branch ca2-soundness on readwidth b970dda; cryptographer)
|
||||
|
||||
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, other agents' builds and the readwidth measurements running beside (the Mac measure lock was free during the GPU runs; nothing here is a hash-rate figure). Write-up `docs/analysis/scratch-soundness.md`; tests `igneum-pow/tests/scratch.rs`; harness `proto-metal/packbench` built from this branch (`--batch-base` added) into the session scratchpad with `swiftc -O -target arm64-apple-macos11 -framework Metal`.
|
||||
|
||||
CPU, `with-lock.sh build nice -n 19 ~/.cargo/bin/cargo test -j4 --test scratch -- --nocapture` (3.6 s): 7 of 7 pass. Stats, 6 classes x 3 seeds x 2^11 units, every read-modify-write traced (3.1 to 12.6 million per class): written-word bias within 6 sigma (worst 3.63); re-hit rate measured against the uniform birthday rate 12.58 vs 10.91 percent (scr2k32), 21.47 vs 20.83 (scr4k32), 36.99 vs 36.50 (scr8k32), 3.84 vs 2.88 (scr2k128), 6.25 vs 5.82 (scr4k128), 11.89 vs 11.37 (scr8k128); slot histogram non-uniform (hottest slot 1.39x to 5.10x the mean: the slot is a register's low bits); deepest chain 5 to 9. Edge: 7 hand-built programs x 2 geometries x 4 bases against an independent hand model, 56 of 56, and 56 of 56 mismatches with the hand model's rewrite words swapped. Static scratch check: 42 of 42 emitted kernels of the 7 scr packs (regenerated byte for byte from program.json first), 6 deliberate breaks caught. Fuzz: 200 generated scratch programs, contract and acceptance on every instruction, 800 units; `IGNEUM_SCRATCH_PACKS_OUT` wrote 214 packs (57 s, three memory-hard caches). The crate's other tests: 33 of 34 lib tests pass; `verify::tests::fold_and_wide_fetch` fails on the readwidth tip itself (`verify.rs:508`, `k as u32 * 0x9E37_79B1` overflows under the test profile; not touched here).
|
||||
|
||||
Metal, `with-lock.sh run <script>`, scripts `gpu-a.sh` and `gpu-b.sh` in the session scratchpad (one `packbench` call per line):
|
||||
|
||||
| Run | Command shape | Result |
|
||||
|---|---|---|
|
||||
| scr4k32 standard pack, timing | `packbench --pack proto-cuda/packs-readwidth/scr4k32 --batches 1 --batch-log2 24 --warps 2048` | 3/3 standalone, 3/3 in batch, fingerprint `3d1af881bd978fb9`, 1.8 s wall for the run |
|
||||
| warp-count independence | same pack, `--batch-log2 12 --warps 1`, `2`, `128` | fingerprint `8c07620f4d9adefd` at all three |
|
||||
| wrap inside the launch | `--batch-log2 9 --batch-base 4294967040 --warps 1`, `4` | fingerprint `8e9e233234d3a297` at both, the base-0 vector inside the window after the wrap 1/1 |
|
||||
| broken tag (`tag = salt`) on a copy of scr4k32, standard vectors | `--batch-log2 24 --warps 2048` | standalone 3/3, in batch 2/3 (base 1,000,000, warp 530's 16th unit, caught), overall FAIL |
|
||||
| broken tag, 2 units on 1 warp, standard vectors | `--batch-log2 6 --warps 1` | 3/3, 1/1, PASS: missed, the standard vectors have no base 32 |
|
||||
| 14 edge packs, run A | `--batch-log2 6 --warps 1 --batches 1` | 14/14 PASS, 4/4 standalone and 2/2 in batch each (bases 0 and 32 on one arena) |
|
||||
| 14 edge packs, run B | `--batch-log2 9 --warps 1 --batch-base 4294967040` | 14/14 PASS, 4/4 and 3/3 each (16 units on one arena, the wrap inside) |
|
||||
| broken tag on edge-slot0 at 32 and 128 KiB | `--batch-log2 6 --warps 1` | 4/4 standalone, 1/2 in batch, FAIL at both (the second unit read the first's slot 0) |
|
||||
| broken lazy fill (`m_` all ones) on edge-slot0 at 32 KiB | same | 0/4, 0/2, FAIL |
|
||||
| 200 fuzz packs | `--batch-log2 9 --warps 2 --batch-base 4294967040 --batches 1` each | 200/200 PASS, 800/800 standalone units (25,600 hashes), 400/400 in the window; 91 s wall for the 200 runs (20:16:02 to 20:17:33 UTC) |
|
||||
|
||||
Totals: 228 of 228 PASS where expected, 3 of 3 FAIL where built in. Reading: the one-warp CPU simulation is exact on Metal under the hosts' present tag policy; the analysis names the host contract (zero the arena at allocation and at the tag counter's wrap, tags from 1, groups a multiple of warps) that turns that into a guarantee, and finds layer 3 does not move the named chip (section 3.4 of the write-up: 2.4x at every share under the cap).
|
||||
|
|
|
|||
|
|
@ -51,6 +51,22 @@ pub struct ScratchModel {
|
|||
data: Vec<[u32; 3]>,
|
||||
pub reads: usize,
|
||||
pub writes: usize,
|
||||
/// Soundness tests (`tests/scratch.rs`, `docs/analysis/scratch-soundness.md`): when `Some`, every
|
||||
/// read-modify-write is appended as it happened. `None` on every verification path.
|
||||
pub trace: Option<Vec<ScratchEvent>>,
|
||||
}
|
||||
|
||||
/// One scratch read-modify-write as the interpreter saw it (variant 5 soundness tests).
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct ScratchEvent {
|
||||
pub lane: u8,
|
||||
pub slot: u32,
|
||||
/// The slot had been written earlier in this unit (a re-hit): the words read were a rewrite, not the fill.
|
||||
pub hit: bool,
|
||||
pub read: [u32; 3],
|
||||
/// The fold result, the new value of `dst`.
|
||||
pub x: u32,
|
||||
pub written: [u32; 3],
|
||||
}
|
||||
|
||||
impl ScratchModel {
|
||||
|
|
@ -61,6 +77,7 @@ impl ScratchModel {
|
|||
data: vec![[0; 3]; LANES * slots_per_lane],
|
||||
reads: 0,
|
||||
writes: 0,
|
||||
trace: None,
|
||||
}
|
||||
}
|
||||
/// Read slot `slot` of `lane`, then rewrite it from the fold result `x`. Returns the three words read.
|
||||
|
|
@ -77,7 +94,11 @@ impl ScratchModel {
|
|||
]
|
||||
};
|
||||
let x = fold_words(dst, &w);
|
||||
self.data[i] = scratch_rewrite(x, &w);
|
||||
let out = scratch_rewrite(x, &w);
|
||||
if let Some(t) = self.trace.as_mut() {
|
||||
t.push(ScratchEvent { lane: lane as u8, slot, hit: self.written[i], read: w, x, written: out });
|
||||
}
|
||||
self.data[i] = out;
|
||||
self.written[i] = true;
|
||||
self.reads += 1;
|
||||
self.writes += 1;
|
||||
|
|
@ -243,6 +264,19 @@ pub fn interpret_warp(program: &Program, base_nonce: u32, ds: &DatasetSource) ->
|
|||
/// [`interpret_warp`] with explicit init words `I` (section 1.6 of the spec). The packs use `I = program.seed`;
|
||||
/// a block uses `I = bind::block_init_words(H, nonce)`.
|
||||
pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, ds: &DatasetSource) -> WarpResult {
|
||||
interpret_warp_scratch(program, seed, base_nonce, ds, false).0
|
||||
}
|
||||
|
||||
/// [`interpret_warp_init`] that also returns every scratch read-modify-write of the unit in execution order
|
||||
/// (lane-minor within an instruction, as the interpreter runs them) when `trace` is set; empty otherwise and for
|
||||
/// a class without a scratch. For the soundness tests of variant 5 only.
|
||||
pub fn interpret_warp_scratch(
|
||||
program: &Program,
|
||||
seed: &[u32; 8],
|
||||
base_nonce: u32,
|
||||
ds: &DatasetSource,
|
||||
trace: bool,
|
||||
) -> (WarpResult, Vec<ScratchEvent>) {
|
||||
let mask = ds.mask;
|
||||
let mut r = [[0u32; LANES]; 8];
|
||||
for lane in 0..LANES {
|
||||
|
|
@ -258,6 +292,11 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
|
|||
let mut idx = [0u32; LANES];
|
||||
let mut val = [0u32; LANES];
|
||||
let mut scratch = if program.has_scratch() { Some(ScratchModel::new(program.class.scratch_slots_per_lane())) } else { None };
|
||||
if trace {
|
||||
if let Some(m) = scratch.as_mut() {
|
||||
m.trace = Some(Vec::new());
|
||||
}
|
||||
}
|
||||
let slot_mask = program.class.scratch_slot_mask();
|
||||
for _ in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
|
|
@ -279,7 +318,8 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
|
|||
let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
|
||||
hashes[lane] = ((hi as u64) << 32) | lo as u64;
|
||||
}
|
||||
WarpResult { hashes, items_derived }
|
||||
let events = scratch.and_then(|m| m.trace).unwrap_or_default();
|
||||
(WarpResult { hashes, items_derived }, events)
|
||||
}
|
||||
|
||||
#[inline(always)]
|
||||
|
|
|
|||
765
igneum-pow/tests/scratch.rs
Normal file
765
igneum-pow/tests/scratch.rs
Normal file
|
|
@ -0,0 +1,765 @@
|
|||
//! Soundness tests of layer 3 of `docs/plans/counter-asic-2.md`: the per-warp scratch with read-modify-writes
|
||||
//! (variant 5 of the read-width experiment, `LoadClass::scratch(k, kb)`). Analysis and results:
|
||||
//! `docs/analysis/scratch-soundness.md`. Every test is parametric over the class's slot count
|
||||
//! (`scratch_slots_per_lane()`), so the 32 and 128 KiB geometries and any later one run the same checks.
|
||||
//!
|
||||
//! What runs under plain `cargo test`:
|
||||
//! 1. `rewrite_is_a_bijection_of_the_fold_value`, `fill_is_a_bijection_of_the_nonce`: the written words as
|
||||
//! functions (question 1).
|
||||
//! 2. `written_words_unbiased_and_rehit_rates`: bit bias of every written word over 2^11 units x 3 seeds per class
|
||||
//! (the TESTS.md section 3 shape), and the measured slot re-hit rate against the birthday formula (question 2).
|
||||
//! 3. `edge_programs_match_the_hand_model`: hand-built programs that drive every read-modify-write of a hash to
|
||||
//! slot 0, slot MASK, through out-of-range registers, to one slot per lane, alternating two slots, and 16
|
||||
//! read-modify-writes per iteration on one slot; the interpreter against an independent hand model, and the
|
||||
//! hand model shown to have teeth (question 3, CPU half).
|
||||
//! 4. `scr_packs_regenerate_and_pass_the_static_scratch_check`: every emitted kernel of every scr pack under
|
||||
//! `proto-cuda/packs-readwidth` regenerates from its program.json and passes the static scratch-mask check;
|
||||
//! the check is shown to fail on four deliberate breaks (question 4).
|
||||
//! 5. `fuzz_scr_programs_cpu`: 200 generated scratch programs over the six classes, generator contract on every
|
||||
//! instruction, 4 units each at base nonces across the 32-bit range including the wrap; with
|
||||
//! `IGNEUM_SCRATCH_PACKS_OUT=<dir>` it also writes the packs (and the edge packs) for the Metal runs of
|
||||
//! `proto-metal/packbench` (question 3 GPU half, question 4, `TESTS.md` section 9 shape).
|
||||
|
||||
use igneum_pow::emit::{
|
||||
cuda_kernel, cuda_kernel_bound, export_pack, metal_program, metal_program_bound, opencl_kernel,
|
||||
opencl_kernel_bound, vectors_json, LoadSource,
|
||||
};
|
||||
use igneum_pow::generator::{
|
||||
generate_class, generate_from_seed_bytes_class, Instr, LoadClass, Op, Program, GENERATOR_VERSION, INSTR_COUNT,
|
||||
ITERATIONS, LANES,
|
||||
};
|
||||
use igneum_pow::seed::{seed_words_from_bytes, SplitMix64};
|
||||
use igneum_pow::verify::{
|
||||
fold_words, interpret_warp_scratch, scratch_fill, scratch_rewrite, splitmix32, DatasetMode, DatasetSource,
|
||||
Epoch, ScratchEvent, FOLD_MUL, FOLD_ROT,
|
||||
};
|
||||
use serde_json::Value;
|
||||
use std::collections::HashMap;
|
||||
use std::path::PathBuf;
|
||||
|
||||
/// The classes under study: the two capped geometries (32 and 128 KiB per warp: 64 and 256 slots per lane) at the
|
||||
/// RMW shares the readwidth branch measures.
|
||||
const CLASSES: [&str; 6] = ["scr2k32", "scr4k32", "scr8k32", "scr2k128", "scr4k128", "scr8k128"];
|
||||
|
||||
fn class(name: &str) -> LoadClass {
|
||||
LoadClass::parse(name).unwrap_or_else(|| panic!("class {name}"))
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
// 1. The written words as functions (question 1)
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
|
||||
/// For a fixed slot content `w`, each of the three rewritten words is a bijection of the fold value `x`
|
||||
/// (`x ^ w1`, `rotl(x, 7) ^ w2`, `x + w0`), so the rewrite is injective in `x` and a uniform `x` gives a uniform
|
||||
/// word in every position. Checked over 2^16 consecutive `x` for 16 random `w`.
|
||||
#[test]
|
||||
fn rewrite_is_a_bijection_of_the_fold_value() {
|
||||
let mut rng = SplitMix64::new(0x7363_7261_7463_6801);
|
||||
for _ in 0..16 {
|
||||
let w = [rng.next() as u32, rng.next() as u32, rng.next() as u32];
|
||||
let x0 = rng.next() as u32;
|
||||
let mut seen = [vec![false; 1 << 16], vec![false; 1 << 16], vec![false; 1 << 16]];
|
||||
for i in 0..(1u32 << 16) {
|
||||
let x = x0.wrapping_add(i);
|
||||
let out = scratch_rewrite(x, &w);
|
||||
for j in 0..3 {
|
||||
// a bijection of x maps 2^16 consecutive x to 2^16 distinct words; the low 16 bits alone are
|
||||
// distinct for the xor words (x ^ c) and for the add word (x + c), since both act on the low 16
|
||||
// bits as bijections of the low 16 bits of x; the rotl word is checked on its rotated-back bits
|
||||
let key = if j == 1 { out[j].rotate_right(7) & 0xffff } else { out[j] & 0xffff };
|
||||
assert!(!seen[j][key as usize], "word {j} repeats inside 2^16 consecutive x");
|
||||
seen[j][key as usize] = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
// The rewrite inverts: from the old content and any ONE written word the fold value is recovered, so a
|
||||
// rewritten slot carries exactly 32 bits of new state (the point of question 2's arithmetic).
|
||||
let w = [0x1234_5678, 0x9abc_def0, 0x0fed_cba9];
|
||||
let x = 0xdead_beef;
|
||||
let out = scratch_rewrite(x, &w);
|
||||
assert_eq!(out[0] ^ w[1], x);
|
||||
assert_eq!((out[1] ^ w[2]).rotate_right(7), x);
|
||||
assert_eq!(out[2].wrapping_sub(w[0]), x);
|
||||
}
|
||||
|
||||
/// For a fixed (seed, slot, j) the fill is a bijection of the lane nonce: `splitmix32` is a bijection of its
|
||||
/// 32-bit input and the input `((base + lane) ^ s) + c` is a bijection of `base + lane`. Over 2^16 consecutive
|
||||
/// nonces no fill word repeats, for 8 slots x 3 words.
|
||||
#[test]
|
||||
fn fill_is_a_bijection_of_the_nonce() {
|
||||
let seed = seed_words_from_bytes(b"igneum-genesis");
|
||||
for slot in [0u32, 1, 63, 64, 255, 1023, 2047] {
|
||||
for j in 0..3u32 {
|
||||
let mut words: Vec<u32> = (0..(1u32 << 16)).map(|n| scratch_fill(&seed, n, 0, slot, j)).collect();
|
||||
words.sort_unstable();
|
||||
words.dedup();
|
||||
assert_eq!(words.len(), 1 << 16, "slot {slot} word {j}: fill words of 2^16 consecutive nonces are distinct");
|
||||
}
|
||||
}
|
||||
// base + lane is the lane nonce: the fill of lane l at base b is the fill of lane 0 at base b + l
|
||||
assert_eq!(scratch_fill(&seed, 0x1000, 7, 5, 2), scratch_fill(&seed, 0x1007, 0, 5, 2));
|
||||
// and it wraps with the nonce: base 0xffffffe0, lane 31 is nonce 0xffffffff; lane 32 would be nonce 0
|
||||
assert_eq!(scratch_fill(&seed, 0xffff_ffe0, 32, 5, 2), scratch_fill(&seed, 0, 0, 5, 2));
|
||||
// the three word positions of one slot and nonce are three different permutation outputs
|
||||
let f: Vec<u32> = (0..3).map(|j| scratch_fill(&seed, 12345, 7, 17, j)).collect();
|
||||
assert!(f[0] != f[1] && f[1] != f[2] && f[0] != f[2]);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
// 2. Uniformity of the written words and the slot re-hit rate (questions 1 and 2)
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
|
||||
/// Birthday arithmetic: the expected number of distinct slots after `n` uniform draws from `s` slots.
|
||||
fn expected_distinct(s: usize, n: usize) -> f64 {
|
||||
let s = s as f64;
|
||||
s * (1.0 - (1.0 - 1.0 / s).powi(n as i32))
|
||||
}
|
||||
|
||||
struct ClassStats {
|
||||
units: usize,
|
||||
events: usize,
|
||||
hits: usize,
|
||||
/// ones count per bit of the written words, 3 x 32
|
||||
ones: [[u64; 32]; 3],
|
||||
/// ones count per bit of written XOR read (the change the rewrite makes to the slot)
|
||||
delta_ones: [[u64; 32]; 3],
|
||||
/// re-hit depth histogram: how many earlier RMWs the slot had seen in this unit (0 = first touch)
|
||||
depth: Vec<usize>,
|
||||
max_depth: usize,
|
||||
/// how often each slot index was addressed (the slot comes from a register's low bits)
|
||||
slot_hist: Vec<u64>,
|
||||
}
|
||||
|
||||
fn class_stats(name: &str, seeds: &[&str], units_per_seed: usize) -> ClassStats {
|
||||
let c = class(name);
|
||||
let mut st = ClassStats {
|
||||
units: 0,
|
||||
events: 0,
|
||||
hits: 0,
|
||||
ones: [[0; 32]; 3],
|
||||
delta_ones: [[0; 32]; 3],
|
||||
depth: vec![0; 256],
|
||||
max_depth: 0,
|
||||
slot_hist: vec![0; c.scratch_slots_per_lane()],
|
||||
};
|
||||
let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 28);
|
||||
for seed in seeds {
|
||||
let p = generate_class(seed, c);
|
||||
assert_eq!(p.scratch_ops_per_hash(), c.scratch_slots() * ITERATIONS);
|
||||
for u in 0..units_per_seed {
|
||||
let base = (u as u32).wrapping_mul(32).wrapping_add(0x4000_0000);
|
||||
let (_, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true);
|
||||
assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES);
|
||||
let mut count: HashMap<(u8, u32), usize> = HashMap::new();
|
||||
for e in &ev {
|
||||
assert!(e.slot < c.scratch_slots_per_lane() as u32, "slot inside the lane's scratch");
|
||||
let d = count.entry((e.lane, e.slot)).or_insert(0);
|
||||
assert_eq!(e.hit, *d > 0, "hit flag agrees with the unit's own history");
|
||||
assert_eq!(e.written, scratch_rewrite(e.x, &e.read));
|
||||
if !e.hit {
|
||||
let fill = [
|
||||
scratch_fill(&p.seed, base, e.lane as u32, e.slot, 0),
|
||||
scratch_fill(&p.seed, base, e.lane as u32, e.slot, 1),
|
||||
scratch_fill(&p.seed, base, e.lane as u32, e.slot, 2),
|
||||
];
|
||||
assert_eq!(e.read, fill, "a first touch reads the fill");
|
||||
}
|
||||
st.depth[(*d).min(255)] += 1;
|
||||
st.max_depth = st.max_depth.max(*d);
|
||||
st.slot_hist[e.slot as usize] += 1;
|
||||
*d += 1;
|
||||
st.events += 1;
|
||||
st.hits += e.hit as usize;
|
||||
for j in 0..3 {
|
||||
for b in 0..32 {
|
||||
st.ones[j][b] += ((e.written[j] >> b) & 1) as u64;
|
||||
st.delta_ones[j][b] += (((e.written[j] ^ e.read[j]) >> b) & 1) as u64;
|
||||
}
|
||||
}
|
||||
}
|
||||
st.units += 1;
|
||||
}
|
||||
}
|
||||
st
|
||||
}
|
||||
|
||||
/// Bit bias of every written word (and of the change each rewrite makes) within 6 sigma of a fair coin, over
|
||||
/// 3 seeds x 2^11 units per class (131,072 hashes per seed set); the slot re-hit rate against the birthday
|
||||
/// formula within 3 percent relative. The table printed here is the one in the analysis.
|
||||
#[test]
|
||||
fn written_words_unbiased_and_rehit_rates() {
|
||||
let seeds = ["igneum-genesis", "igneum-genesis/stats1", "igneum-genesis/stats2"];
|
||||
let units = 1usize << 11;
|
||||
println!("class | slots/lane | RMW/hash | events | re-hits | re-hit % | birthday % | slot chi2 z (spread) | max depth | max bias sigma | max delta bias sigma");
|
||||
for name in CLASSES {
|
||||
let c = class(name);
|
||||
let st = class_stats(name, &seeds, units);
|
||||
let n = st.events as f64;
|
||||
let sigma = (n / 4.0).sqrt();
|
||||
let mut worst = 0.0f64;
|
||||
let mut worst_delta = 0.0f64;
|
||||
for j in 0..3 {
|
||||
for b in 0..32 {
|
||||
let z = (st.ones[j][b] as f64 - n / 2.0).abs() / sigma;
|
||||
let zd = (st.delta_ones[j][b] as f64 - n / 2.0).abs() / sigma;
|
||||
assert!(z <= 6.0, "{name}: written word {j} bit {b} biased: {z:.2} sigma");
|
||||
assert!(zd <= 6.0, "{name}: rewrite delta word {j} bit {b} biased: {zd:.2} sigma");
|
||||
worst = worst.max(z);
|
||||
worst_delta = worst_delta.max(zd);
|
||||
}
|
||||
}
|
||||
let per_lane_hash = c.scratch_slots() * ITERATIONS;
|
||||
let s = c.scratch_slots_per_lane();
|
||||
let exp_hits = per_lane_hash as f64 - expected_distinct(s, per_lane_hash);
|
||||
let exp_pct = 100.0 * exp_hits / per_lane_hash as f64;
|
||||
let got_pct = 100.0 * st.hits as f64 / st.events as f64;
|
||||
// chi-square of the slot histogram against uniform (df = s - 1): the slot is a register's low bits, and
|
||||
// the measured re-hit rate runs above the uniform birthday rate (the finding of the analysis, question 2)
|
||||
let expect_per_slot = n / s as f64;
|
||||
let chi2: f64 = st.slot_hist.iter().map(|&h| (h as f64 - expect_per_slot).powi(2) / expect_per_slot).sum();
|
||||
let chi2_z = (chi2 - (s as f64 - 1.0)) / (2.0 * (s as f64 - 1.0)).sqrt();
|
||||
let hot = *st.slot_hist.iter().max().unwrap() as f64 / expect_per_slot;
|
||||
let cold = *st.slot_hist.iter().min().unwrap() as f64 / expect_per_slot;
|
||||
println!(
|
||||
"{name} | {s} | {per_lane_hash} | {} | {} | {got_pct:.2} | {exp_pct:.2} | {chi2_z:.1} (hottest slot {hot:.2}x, coldest {cold:.2}x) | {} | {worst:.2} | {worst_delta:.2}",
|
||||
st.events, st.hits, st.max_depth
|
||||
);
|
||||
// a regression band, not a uniformity claim: the rate sits between the uniform birthday rate and twice it
|
||||
assert!(
|
||||
got_pct >= 0.9 * exp_pct && got_pct <= 2.0 * exp_pct,
|
||||
"{name}: re-hit rate {got_pct:.2}% against birthday {exp_pct:.2}%"
|
||||
);
|
||||
// depth histogram: the number of earlier RMWs a re-hit slot had seen in the unit
|
||||
let shown: Vec<String> = st.depth.iter().take(st.max_depth + 1).enumerate().map(|(d, n)| format!("{d}:{n}")).collect();
|
||||
println!(" depth histogram {}", shown.join(" "));
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
// 3. Hand-built edge programs against an independent hand model (question 3, CPU half)
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
|
||||
fn ins(op: Op, dst: u8, src: u8) -> Instr {
|
||||
Instr { op, dst, src, src2: 0, imm: 0, imm2: 0, rot: 1, bit: 0, mask: 1, width: 1 }
|
||||
}
|
||||
fn add_imm(dst: u8, src: u8, imm: u32) -> Instr {
|
||||
Instr { op: Op::Add, dst, src, src2: 0, imm, imm2: imm, rot: 1, bit: 0, mask: 1, width: 1 }
|
||||
}
|
||||
|
||||
/// A hand-built program of class `c` named `name` (its seed is the name, so its fill words and init words are
|
||||
/// its own). These bypass the generator and the acceptance rule, like `TESTS.md` section 2; `sub r, r` zeroes a
|
||||
/// register as the Swift edge set does.
|
||||
fn edge(name: &str, c: LoadClass, instrs: Vec<Instr>) -> Program {
|
||||
let seed_string = format!("igneum-scratch-edge/{name}");
|
||||
let seed_bytes = seed_string.as_bytes().to_vec();
|
||||
let k = instrs.iter().filter(|i| i.op == Op::Scratch).count();
|
||||
assert_eq!(k, c.scratch_slots(), "{name}: the class carries the program's scratch count");
|
||||
Program {
|
||||
seed: seed_words_from_bytes(&seed_bytes),
|
||||
seed_string,
|
||||
seed_bytes,
|
||||
generator: GENERATOR_VERSION,
|
||||
attempt: 0,
|
||||
class: c,
|
||||
instrs,
|
||||
}
|
||||
}
|
||||
|
||||
/// The edge set for a scratch of `kb` KiB per warp. Each entry: (name, what it drives, program).
|
||||
fn edge_programs(kb: u8) -> Vec<(String, &'static str, Program)> {
|
||||
let m = LoadClass::scratch(1, kb).scratch_slot_mask();
|
||||
let dsts = [2u8, 3, 4, 5, 6, 7, 0, 2, 3, 4, 5, 6, 7, 0, 2, 3];
|
||||
let scr = |n: usize, src: u8| -> Vec<Instr> { (0..n).map(|i| ins(Op::Scratch, dsts[i], src)).collect() };
|
||||
let mut v = Vec::new();
|
||||
// every RMW of the hash to slot 0 through a zero register: 64 dependent RMWs on one slot per lane
|
||||
let mut p = vec![ins(Op::Sub, 1, 1)];
|
||||
p.extend(scr(8, 1));
|
||||
v.push(("slot0".to_string(), "r1 = 0: every RMW to slot 0", edge(&format!("slot0/k{kb}"), LoadClass::scratch(8, kb), p)));
|
||||
// slot MASK through the in-range register MASK
|
||||
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m)];
|
||||
p.extend(scr(8, 1));
|
||||
v.push(("slotmask".to_string(), "r1 = MASK: every RMW to the last slot", edge(&format!("slotmask/k{kb}"), LoadClass::scratch(8, kb), p)));
|
||||
// slot MASK through the out-of-range register 0xffffffff
|
||||
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, 1), ins(Op::Sub, 1, 2)];
|
||||
p.extend(scr(8, 1));
|
||||
v.push(("ones".to_string(), "r1 = 0xffffffff: masked to the last slot", edge(&format!("ones/k{kb}"), LoadClass::scratch(8, kb), p)));
|
||||
// slot 0 through the out-of-range register MASK + 1
|
||||
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m.wrapping_add(1))];
|
||||
p.extend(scr(8, 1));
|
||||
v.push(("maskplus1".to_string(), "r1 = MASK + 1: masked to slot 0", edge(&format!("maskplus1/k{kb}"), LoadClass::scratch(8, kb), p)));
|
||||
// 16 RMWs per iteration on slot 0: 128 dependent RMWs on one slot per lane per hash
|
||||
let mut p = vec![ins(Op::Sub, 1, 1)];
|
||||
p.extend(scr(16, 1));
|
||||
v.push(("sixteen".to_string(), "16 RMWs per iteration on slot 0", edge(&format!("sixteen/k{kb}"), LoadClass::scratch(16, kb), p)));
|
||||
// one slot per lane from the init words: lanes with equal slots would show any cross-lane aliasing
|
||||
// (r5 is the slot register and is never a destination here)
|
||||
let p: Vec<Instr> = [0u8, 1, 2, 3, 4, 6, 7, 0].iter().map(|&d| ins(Op::Scratch, d, 5)).collect();
|
||||
v.push(("lanevar".to_string(), "r5 never written: one init-dependent slot per lane", edge(&format!("lanevar/k{kb}"), LoadClass::scratch(8, kb), p)));
|
||||
// alternating slot 0 and slot MASK inside one iteration
|
||||
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, m)];
|
||||
for (i, &d) in [3u8, 4, 5, 6, 7, 0, 3, 4].iter().enumerate() {
|
||||
// r1 and r2 hold the two slots and are never destinations
|
||||
p.push(ins(Op::Scratch, d, if i % 2 == 0 { 1 } else { 2 }));
|
||||
}
|
||||
v.push(("twoslots".to_string(), "slot 0 and slot MASK alternating", edge(&format!("twoslots/k{kb}"), LoadClass::scratch(8, kb), p)));
|
||||
v
|
||||
}
|
||||
|
||||
/// The hand model: a second, minimal interpreter for the ops the edge programs use (sub, add, scratch), with its
|
||||
/// own slot store keyed by (lane, slot). `mutate` swaps the rewrite's words to show the comparison has teeth.
|
||||
fn hand_model(p: &Program, base: u32, mutate: bool) -> [u64; 32] {
|
||||
let seed = &p.seed;
|
||||
let m = p.class.scratch_slot_mask();
|
||||
let mut r = [[0u32; LANES]; 8];
|
||||
for lane in 0..LANES {
|
||||
let nonce = base.wrapping_add(lane as u32);
|
||||
for i in 0..8 {
|
||||
let mut x = nonce ^ seed[i];
|
||||
x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1));
|
||||
x = splitmix32(x);
|
||||
r[i][lane] = x ^ seed[(i + 1) & 7];
|
||||
}
|
||||
}
|
||||
let mut store: HashMap<(usize, u32), [u32; 3]> = HashMap::new();
|
||||
for _ in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
for ins in &p.instrs {
|
||||
let (d, a) = (ins.dst as usize, ins.src as usize);
|
||||
match ins.op {
|
||||
Op::Sub => {
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = r[d][lane].wrapping_sub(r[a][lane]);
|
||||
}
|
||||
}
|
||||
Op::Add => {
|
||||
for lane in 0..LANES {
|
||||
let c = if (sel[lane] >> ins.bit) & 1 != 0 { ins.imm2 } else { ins.imm };
|
||||
r[d][lane] = r[d][lane].wrapping_add(r[a][lane]).wrapping_add(c);
|
||||
}
|
||||
}
|
||||
Op::Scratch => {
|
||||
for lane in 0..LANES {
|
||||
let slot = r[a][lane] & m;
|
||||
let w = *store.entry((lane, slot)).or_insert_with(|| {
|
||||
let mut f = [0u32; 3];
|
||||
for j in 0..3u32 {
|
||||
// the fill, written out in full rather than through verify::scratch_fill
|
||||
let n = base.wrapping_add(lane as u32);
|
||||
f[j as usize] = splitmix32(
|
||||
(n ^ seed[j as usize])
|
||||
.wrapping_add(slot.wrapping_mul(0x9E37_79B1))
|
||||
.wrapping_add((j + 1).wrapping_mul(0x85EB_CA77)),
|
||||
);
|
||||
}
|
||||
f
|
||||
});
|
||||
let mut x = r[d][lane] ^ w[0];
|
||||
x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[1];
|
||||
x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[2];
|
||||
r[d][lane] = x;
|
||||
let out = if mutate {
|
||||
[x.rotate_left(7) ^ w[2], x ^ w[1], x.wrapping_add(w[0])]
|
||||
} else {
|
||||
[x ^ w[1], x.rotate_left(7) ^ w[2], x.wrapping_add(w[0])]
|
||||
};
|
||||
store.insert((lane, slot), out);
|
||||
}
|
||||
}
|
||||
other => panic!("the hand model does not implement {other:?}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
let mut out = [0u64; 32];
|
||||
for lane in 0..LANES {
|
||||
let lo = r[0][lane] ^ r[1][lane].rotate_left(7) ^ r[2][lane].rotate_left(14) ^ r[3][lane].rotate_left(21);
|
||||
let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
|
||||
out[lane] = ((hi as u64) << 32) | lo as u64;
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// The four unit bases of every edge vector: 0 and 32 (two consecutive units, the pair a one-warp persistent
|
||||
/// launch runs on one arena), a unit straddling 2^31, and the unit that wraps past 2^32.
|
||||
const EDGE_BASES: [u32; 4] = [0, 32, 0x7fff_fff0, 0xffff_ffe0];
|
||||
|
||||
#[test]
|
||||
fn edge_programs_match_the_hand_model() {
|
||||
let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 24);
|
||||
let mut cases = 0;
|
||||
for kb in [32u8, 128] {
|
||||
for (name, what, p) in edge_programs(kb) {
|
||||
let slots = p.class.scratch_slots_per_lane();
|
||||
for base in EDGE_BASES {
|
||||
let (res, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true);
|
||||
let hand = hand_model(&p, base, false);
|
||||
assert_eq!(res.hashes, hand, "{name} k{kb} base {base:#x}: interpreter against the hand model ({what})");
|
||||
assert_ne!(res.hashes, hand_model(&p, base, true), "{name} k{kb}: the comparison has teeth");
|
||||
// the slots the trace saw are the ones the program was built to drive
|
||||
let slot_set: std::collections::BTreeSet<u32> = ev.iter().map(|e| e.slot).collect();
|
||||
let m = (slots - 1) as u32;
|
||||
match name.as_str() {
|
||||
"slot0" | "maskplus1" | "sixteen" => assert_eq!(slot_set.into_iter().collect::<Vec<_>>(), vec![0]),
|
||||
"slotmask" | "ones" => assert_eq!(slot_set.into_iter().collect::<Vec<_>>(), vec![m]),
|
||||
"twoslots" => assert_eq!(slot_set.into_iter().collect::<Vec<_>>(), vec![0, m]),
|
||||
"lanevar" => {
|
||||
for e in &ev {
|
||||
assert!(e.slot <= m);
|
||||
}
|
||||
}
|
||||
_ => unreachable!(),
|
||||
}
|
||||
// the chain depth on the driven slot: every RMW after the first per lane is a re-hit
|
||||
let per_lane = p.scratch_ops_per_hash();
|
||||
let hits = ev.iter().filter(|e| e.hit).count();
|
||||
let expected_hits = match name.as_str() {
|
||||
"twoslots" => (per_lane - 2) * LANES,
|
||||
_ => (per_lane - 1) * LANES,
|
||||
};
|
||||
assert_eq!(hits, expected_hits, "{name} k{kb}: re-hits");
|
||||
cases += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
assert_eq!(cases, 2 * 7 * 4);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
// 4. The static scratch check over every emitted kernel of every scr pack (question 4)
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum Dialect {
|
||||
Metal,
|
||||
Cuda,
|
||||
OpenCl,
|
||||
}
|
||||
|
||||
/// The static scratch check: every scratch read-modify-write in an emitted kernel has the one masked form the
|
||||
/// emitter writes, the arena is the lane's own `slots x 4` words, the tag is `salt + unit`, and nothing else
|
||||
/// touches the scratch. Like the dataset mask check of `TESTS.md` section 5 and `tests/packs.rs`, a text check:
|
||||
/// the guarantee is that the emitter has one template and it masks.
|
||||
pub fn scratch_text_check(text: &str, dialect: Dialect, k: usize, slots: usize, kernels: usize) -> Result<(), String> {
|
||||
assert!(kernels >= 1);
|
||||
// every count below is per hash kernel; an OpenCL bound file carries igneum_hash and igneum_hash_bound
|
||||
let k = k * kernels;
|
||||
assert!(slots.is_power_of_two() && slots >= 1);
|
||||
let mask = (slots - 1) as u32;
|
||||
let wpl = slots * 4;
|
||||
let (u, load, store, ptr) = match dialect {
|
||||
Dialect::Metal => ("uint", "uint4 v_ = *(device const uint4*)(arena + s_ * 4u);", "*(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "device uint* arena"),
|
||||
Dialect::Cuda => ("uint32_t", "uint4 v_ = *(const uint4*)(arena + s_ * 4u);", "*(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "uint32_t* arena"),
|
||||
Dialect::OpenCl => ("uint", "uint4 v_ = vload4(s_, arena);", "vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); }", "__global uint* arena"),
|
||||
};
|
||||
let count = |needle: &str| text.matches(needle).count();
|
||||
let mut errs = Vec::new();
|
||||
let mut expect = |what: &str, got: usize, want: usize| {
|
||||
if got != want {
|
||||
errs.push(format!("{what}: {got}, expected {want}"));
|
||||
}
|
||||
};
|
||||
// k slot computations, each masked with exactly the class's mask and immediately followed by the one load form
|
||||
expect("slot definitions `{ u s_ = r`", count(&format!("{{ {u} s_ = r")), k);
|
||||
expect("masked slot followed by the load", count(&format!(" & {mask}u; {load}")), k);
|
||||
expect("stores of the tagged slot", count(store), k);
|
||||
expect("tag compares", count("(v_.x == tag)"), k);
|
||||
expect("fill calls (three per RMW)", count("scr_fill(gbase, lane, s_, "), 3 * k);
|
||||
// the arena: one definition with the class's words per lane, and 2k uses (one load, one store per RMW)
|
||||
expect("arena definition", count(&format!("{ptr} = scratch + ((size_t)warp_ * 32u + lane) * {wpl}u;")), kernels);
|
||||
expect("arena mentions (definition + load + store per RMW)", count("arena"), kernels + 2 * k);
|
||||
expect("tag definition `tag = salt + g_`", count(&format!("{u} tag = salt + g_;")), kernels);
|
||||
expect("direct scratch indexing", count("scratch["), 0);
|
||||
expect("scratch pointer arithmetic outside the arena definition", count("scratch +"), kernels);
|
||||
// no other mask value on a slot: every `s_ = r` line carries the class mask and nothing else carries ` & Nu; uint4 v_`
|
||||
let any_mask_load = count(&format!("u; {load}"));
|
||||
expect("loads preceded by some mask (must all be the class mask)", any_mask_load, k);
|
||||
if errs.is_empty() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(errs.join("; "))
|
||||
}
|
||||
}
|
||||
|
||||
fn packs_rw_dir() -> PathBuf {
|
||||
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-readwidth")
|
||||
}
|
||||
|
||||
fn scr_packs() -> Vec<String> {
|
||||
let mut v: Vec<String> = std::fs::read_dir(packs_rw_dir())
|
||||
.unwrap()
|
||||
.map(|d| d.unwrap().file_name().to_string_lossy().to_string())
|
||||
.filter(|n| n.starts_with("scr"))
|
||||
.collect();
|
||||
v.sort();
|
||||
v
|
||||
}
|
||||
|
||||
fn read_pack(pack: &str, file: &str) -> String {
|
||||
let p = packs_rw_dir().join(pack).join(file);
|
||||
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
|
||||
}
|
||||
|
||||
/// Every scr pack regenerates from its program.json (seed bytes, class, day bytes, size) to the same six kernel
|
||||
/// texts, byte for byte, and every one of those texts passes the static scratch check for the class's k and slot
|
||||
/// count; the check fails on four deliberate breaks of a copy of the Metal text (mask dropped, mask changed, arena
|
||||
/// stride changed, a stray scratch access) and on the OpenCL and CUDA twins of the first.
|
||||
#[test]
|
||||
fn scr_packs_regenerate_and_pass_the_static_scratch_check() {
|
||||
let packs = scr_packs();
|
||||
assert!(packs.len() >= 6, "the scr packs: {packs:?}");
|
||||
let mut checked = 0;
|
||||
let mut sample_metal = String::new();
|
||||
let mut sample_cl = String::new();
|
||||
let mut sample_cu = String::new();
|
||||
let mut sample_k = 0;
|
||||
let mut sample_slots = 0;
|
||||
for pack in &packs {
|
||||
let j: Value = serde_json::from_str(&read_pack(pack, "program.json")).unwrap();
|
||||
let name = j["load_class"].as_str().unwrap();
|
||||
let c = class(name);
|
||||
assert_eq!(&format!("{name}"), pack, "pack directory named after its class");
|
||||
let seed = j["seed"].as_str().unwrap();
|
||||
let seed_bytes = igneum_pow::bind::unhex(j["seed_bytes"].as_str().unwrap()).unwrap();
|
||||
let day_bytes = igneum_pow::bind::unhex(j["dataset"]["day_bytes"].as_str().unwrap()).unwrap();
|
||||
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
|
||||
assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard");
|
||||
let program = generate_from_seed_bytes_class(seed, &seed_bytes, c);
|
||||
assert_eq!(program.class, c);
|
||||
assert_eq!(program.program_id(), u64::from_str_radix(j["program_id"].as_str().unwrap().trim_start_matches("0x"), 16).unwrap());
|
||||
let mut dataset = DatasetSource::from_key(seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2);
|
||||
dataset.key_bytes = day_bytes;
|
||||
let e = Epoch { program, dataset };
|
||||
let p = &e.program;
|
||||
let mp = e.dataset.memhard().map(|m| &m.params);
|
||||
let k = c.scratch_slots();
|
||||
let slots = c.scratch_slots_per_lane();
|
||||
assert_eq!(p.scratch_ops_per_hash(), k * ITERATIONS);
|
||||
for (file, text, dialect, kernels) in [
|
||||
("program.metal", metal_program(p, log2, LoadSource::Stored), Dialect::Metal, 1),
|
||||
("program_bound.metal", metal_program_bound(p, log2), Dialect::Metal, 1),
|
||||
("kernel.cu", cuda_kernel(p, mp), Dialect::Cuda, 1),
|
||||
("kernel_bound.cu", cuda_kernel_bound(p, mp), Dialect::Cuda, 1),
|
||||
("kernel.cl", opencl_kernel(p, mp), Dialect::OpenCl, 1),
|
||||
// the OpenCL bound file carries igneum_hash and igneum_hash_bound
|
||||
("kernel_bound.cl", opencl_kernel_bound(p, mp), Dialect::OpenCl, 2),
|
||||
] {
|
||||
let on_disk = read_pack(pack, file);
|
||||
assert_eq!(on_disk, text, "{pack}/{file}: the pack is the emitter's text");
|
||||
// scr0 is the persistent control: an arena and a tag, no read-modify-write; the check holds with k = 0
|
||||
scratch_text_check(&on_disk, dialect, k, slots, kernels).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"));
|
||||
checked += 1;
|
||||
}
|
||||
// the vectors of the pack are the CPU's
|
||||
let v: Value = serde_json::from_str(&read_pack(pack, "vectors.json")).unwrap();
|
||||
for w in v["warps"].as_array().unwrap() {
|
||||
let base = w["base_nonce"].as_u64().unwrap() as u32;
|
||||
let got = e.hash_warp(base);
|
||||
for (lane, x) in w["expected"].as_array().unwrap().iter().enumerate() {
|
||||
let want = u64::from_str_radix(x.as_str().unwrap().trim_start_matches("0x"), 16).unwrap();
|
||||
assert_eq!(got[lane], want, "{pack}: base {base} lane {lane}");
|
||||
}
|
||||
}
|
||||
if k == 4 && slots == 64 {
|
||||
sample_metal = read_pack(pack, "program.metal");
|
||||
sample_cl = read_pack(pack, "kernel.cl");
|
||||
sample_cu = read_pack(pack, "kernel.cu");
|
||||
sample_k = k;
|
||||
sample_slots = slots;
|
||||
}
|
||||
}
|
||||
assert_eq!(checked, packs.len() * 6);
|
||||
println!("static scratch check: {checked} kernels over {} scr packs", packs.len());
|
||||
|
||||
// The deliberate breaks (the watcher rule of CLAUDE.md: a check is trusted once it fails on a known-broken
|
||||
// case). Each must be caught; the message names what.
|
||||
assert!(sample_k == 4 && sample_slots == 64, "scr4k32 is in the pack set");
|
||||
let mask = format!(" & {}u; uint4 v_", sample_slots - 1);
|
||||
let broken_mask = sample_metal.replacen(&mask, "; uint4 v_", 1);
|
||||
assert_ne!(broken_mask, sample_metal);
|
||||
let e = scratch_text_check(&broken_mask, Dialect::Metal, 4, 64, 1).unwrap_err();
|
||||
assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
|
||||
println!("break 1 (one mask dropped, Metal): {e}");
|
||||
let wrong_mask = sample_metal.replace(" & 63u;", " & 127u;");
|
||||
let e = scratch_text_check(&wrong_mask, Dialect::Metal, 4, 64, 1).unwrap_err();
|
||||
assert!(e.contains("masked slot followed by the load: 0, expected 4"), "{e}");
|
||||
println!("break 2 (mask 63 -> 127 on every RMW, Metal): {e}");
|
||||
let wrong_stride = sample_metal.replace("* 256u;", "* 128u;");
|
||||
let e = scratch_text_check(&wrong_stride, Dialect::Metal, 4, 64, 1).unwrap_err();
|
||||
assert!(e.contains("arena definition: 0, expected 1"), "{e}");
|
||||
println!("break 3 (arena stride 256 -> 128 words, Metal): {e}");
|
||||
let stray = format!("{sample_metal}\n// stray\n// arena[0] = 0u; scratch[1] = 1u;\n");
|
||||
let e = scratch_text_check(&stray, Dialect::Metal, 4, 64, 1).unwrap_err();
|
||||
assert!(e.contains("arena mentions") && e.contains("direct scratch indexing: 1, expected 0"), "{e}");
|
||||
println!("break 4 (a stray arena and scratch access, Metal): {e}");
|
||||
let e = scratch_text_check(&sample_cl.replacen(" & 63u; uint4 v_ = vload4", "; uint4 v_ = vload4", 1), Dialect::OpenCl, 4, 64, 1).unwrap_err();
|
||||
assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
|
||||
println!("break 5 (one mask dropped, OpenCL): {e}");
|
||||
let e = scratch_text_check(&sample_cu.replacen(" & 63u; uint4 v_ = *(const uint4*)", "; uint4 v_ = *(const uint4*)", 1), Dialect::Cuda, 4, 64, 1).unwrap_err();
|
||||
assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
|
||||
println!("break 6 (one mask dropped, CUDA): {e}");
|
||||
// and the unbroken texts pass under the same calls
|
||||
scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 1).unwrap();
|
||||
scratch_text_check(&sample_cl, Dialect::OpenCl, 4, 64, 1).unwrap();
|
||||
scratch_text_check(&sample_cu, Dialect::Cuda, 4, 64, 1).unwrap();
|
||||
// a wrong slot count, RMW count or kernel count against a right text fails too (the check is tied to the class)
|
||||
assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 256, 1).is_err());
|
||||
assert!(scratch_text_check(&sample_metal, Dialect::Metal, 3, 64, 1).is_err());
|
||||
assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 2).is_err());
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
// 5. The fuzz: 200 generated scratch programs, contract on every instruction, 4 units each across the 32-bit
|
||||
// range including the wrap; with IGNEUM_SCRATCH_PACKS_OUT the packs for the Metal runs (question 3, 4)
|
||||
// ---------------------------------------------------------------------------------------------------------------
|
||||
|
||||
/// Write a pack whose vectors.json carries `bases` (any number of units) instead of the three standard bases.
|
||||
fn write_pack_with_bases(dir: &PathBuf, e: &Epoch, day: &str, bases: &[u32], source: &str) -> Vec<[u64; 32]> {
|
||||
let mut pack = export_pack(e, day, source);
|
||||
let outs: Vec<[u64; 32]> = bases.iter().map(|&b| e.hash_warp(b)).collect();
|
||||
let vj = vectors_json(&e.program, day, e.dataset.log2_words, bases, &outs, &pack.vectors, e.dataset.mask, source, true);
|
||||
for f in pack.files.iter_mut() {
|
||||
if f.0 == "vectors.json" {
|
||||
f.1 = vj.clone();
|
||||
}
|
||||
}
|
||||
pack.write_to(dir).unwrap();
|
||||
outs
|
||||
}
|
||||
|
||||
fn contract(p: &Program) {
|
||||
assert_eq!(p.instrs.len(), INSTR_COUNT);
|
||||
assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count() + p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), 16);
|
||||
assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), p.class.scratch_slots());
|
||||
assert!(p.instrs[0].op != Op::Load && p.instrs[0].op != Op::Scratch, "instruction 0 is never a memory op");
|
||||
for (k, i) in p.instrs.iter().enumerate() {
|
||||
assert!(i.src != i.dst, "#{k}: src == dst");
|
||||
assert!((1..=31).contains(&i.rot), "#{k}: rot {}", i.rot);
|
||||
assert!([1u8, 2, 4, 8, 16].contains(&i.mask), "#{k}: mask {}", i.mask);
|
||||
assert!(i.dst < 8 && i.src < 8 && i.src2 < 8);
|
||||
assert_eq!(i.width, 1, "#{k}: a scratch class reads one-word loads");
|
||||
}
|
||||
assert!(igneum_pow::accept::check(p).is_ok(), "an accepted program");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn fuzz_scr_programs_cpu() {
|
||||
let n: usize = std::env::var("IGNEUM_SCRATCH_FUZZ").ok().and_then(|s| s.parse().ok()).unwrap_or(200);
|
||||
let out = std::env::var("IGNEUM_SCRATCH_PACKS_OUT").ok().map(PathBuf::from);
|
||||
let mut rng = SplitMix64::new(0x6967_6e65_756d_2d73); // "igneum-s"
|
||||
let day = "2026-10-03";
|
||||
let closed = DatasetSource::new(day, DatasetMode::ClosedForm, 28);
|
||||
// memory-hard sources per size, built once each (the cache fill is 0.2 s); only when packs are written
|
||||
let mut mh: HashMap<u32, DatasetSource> = HashMap::new();
|
||||
let mut manifest = String::from("pack\tclass\tlog2\tprogram_id\tscratch_ops_per_hash\tbases\n");
|
||||
let mut per_class: HashMap<String, usize> = HashMap::new();
|
||||
let mut units = 0usize;
|
||||
let mut wraps = 0usize;
|
||||
if let Some(dir) = &out {
|
||||
std::fs::create_dir_all(dir).unwrap();
|
||||
// the edge packs first: 64 MiB datasets (no dataset load in them), the four edge bases
|
||||
for kb in [32u8, 128] {
|
||||
for (name, _what, p) in edge_programs(kb) {
|
||||
let log2 = 24;
|
||||
let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2));
|
||||
let e = Epoch { program: p, dataset: ds };
|
||||
let pack_name = format!("edge-{name}-k{kb}");
|
||||
write_pack_with_bases(&dir.join(&pack_name), &e, day, &EDGE_BASES, "igneum-pow tests/scratch.rs edge");
|
||||
manifest.push_str(&format!(
|
||||
"{pack_name}\t{}\t{log2}\t{:016x}\t{}\t{}\n",
|
||||
e.program.class.name(),
|
||||
e.program.program_id(),
|
||||
e.program.scratch_ops_per_hash(),
|
||||
EDGE_BASES.iter().map(|b| format!("{b}")).collect::<Vec<_>>().join(",")
|
||||
));
|
||||
mh.insert(log2, e.dataset);
|
||||
}
|
||||
}
|
||||
}
|
||||
for i in 0..n {
|
||||
let name = CLASSES[rng.below(CLASSES.len() as u64) as usize];
|
||||
let c = class(name);
|
||||
let seed = format!("igneum-scratch-fuzz/{i}");
|
||||
let p = generate_class(&seed, c);
|
||||
contract(&p);
|
||||
*per_class.entry(name.to_string()).or_insert(0) += 1;
|
||||
// four bases: one inside a 256-nonce batch (in-batch check on the GPU), one straddling 2^31, one in
|
||||
// the last 256 nonces (the unit wraps past 2^32 or ends on it), one uniform
|
||||
let b0 = (rng.below(8) as u32) * 32;
|
||||
let b1 = 0x8000_0000u32.wrapping_sub(256).wrapping_add((rng.below(16) as u32) * 32);
|
||||
let b2 = 0xffff_ff00u32.wrapping_add((rng.below(8) as u32) * 32);
|
||||
let b3 = (rng.next() as u32) & !31;
|
||||
let bases = [b0, b1, b2, b3];
|
||||
// an aligned unit never straddles 2^32 (spec 1.9); the top unit ends on 0xffffffff and the persistent
|
||||
// kernel's unit sequence wraps inside a launch, which the Metal run checks with packbench --batch-base
|
||||
wraps += bases.iter().filter(|&&b| b >= 0xffff_ff00).count();
|
||||
// the CPU: the interpreter is deterministic and every scratch event is inside the lane's slots
|
||||
for &b in &bases {
|
||||
let (r1, ev) = interpret_warp_scratch(&p, &p.seed, b, &closed, true);
|
||||
let r2 = interpret_warp_scratch(&p, &p.seed, b, &closed, false).0;
|
||||
assert_eq!(r1.hashes, r2.hashes);
|
||||
assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES);
|
||||
assert!(ev.iter().all(|e: &ScratchEvent| e.slot < c.scratch_slots_per_lane() as u32));
|
||||
units += 1;
|
||||
}
|
||||
if let Some(dir) = &out {
|
||||
let log2 = [24u32, 26, 28][rng.below(3) as usize];
|
||||
let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2));
|
||||
let e = Epoch { program: p, dataset: ds };
|
||||
let pack_name = format!("fuzz-{i:03}-{name}-l{log2}");
|
||||
write_pack_with_bases(&dir.join(&pack_name), &e, day, &bases, "igneum-pow tests/scratch.rs fuzz");
|
||||
manifest.push_str(&format!(
|
||||
"{pack_name}\t{name}\t{log2}\t{:016x}\t{}\t{}\n",
|
||||
e.program.program_id(),
|
||||
e.program.scratch_ops_per_hash(),
|
||||
bases.iter().map(|b| format!("{b}")).collect::<Vec<_>>().join(",")
|
||||
));
|
||||
mh.insert(log2, e.dataset);
|
||||
} else {
|
||||
let _ = rng.below(3);
|
||||
}
|
||||
}
|
||||
let mut classes: Vec<_> = per_class.iter().collect();
|
||||
classes.sort();
|
||||
println!("fuzz: {n} programs, {units} units on the CPU, {wraps} units in the top 256 nonces, classes {classes:?}");
|
||||
assert_eq!(units, 4 * n);
|
||||
assert_eq!(wraps, n, "every program has a unit in the top 256 nonces");
|
||||
if let Some(dir) = &out {
|
||||
std::fs::write(dir.join("manifest.tsv"), manifest).unwrap();
|
||||
println!("packs written to {}", dir.display());
|
||||
}
|
||||
}
|
||||
|
||||
/// The fold and rewrite, restated: a slot after `d` dependent RMWs holds 96 bits that are a function of the fill
|
||||
/// (3 words, a pure function of nonce, slot and seed) and the `d` fold values; a chip that keeps the `d` fold
|
||||
/// values (32 bits each) instead of the 96-bit slot recomputes the slot in `d` rewrites. This test pins the
|
||||
/// arithmetic the analysis uses (question 2): the replay from the fold values reproduces the slot.
|
||||
#[test]
|
||||
fn slot_is_replayable_from_its_fold_values() {
|
||||
let seed = seed_words_from_bytes(b"igneum-genesis");
|
||||
let (base, lane, slot) = (0x1234_5600u32, 5u32, 17u32);
|
||||
let fill = [scratch_fill(&seed, base, lane, slot, 0), scratch_fill(&seed, base, lane, slot, 1), scratch_fill(&seed, base, lane, slot, 2)];
|
||||
let mut rng = SplitMix64::new(99);
|
||||
let dsts: Vec<u32> = (0..64).map(|_| rng.next() as u32).collect();
|
||||
// the honest sequence: read, fold, rewrite, 64 times
|
||||
let mut w = fill;
|
||||
let mut xs = Vec::new();
|
||||
for &d in &dsts {
|
||||
let x = fold_words(d, &w);
|
||||
xs.push(x);
|
||||
w = scratch_rewrite(x, &w);
|
||||
}
|
||||
// the replay: from the fill and the stored fold values alone
|
||||
let mut w2 = fill;
|
||||
for &x in &xs {
|
||||
w2 = scratch_rewrite(x, &w2);
|
||||
}
|
||||
assert_eq!(w, w2);
|
||||
// and nothing shorter: the fold value at step d depends on the slot content at step d, which depends on
|
||||
// every earlier fold value (drop one and the chain diverges)
|
||||
let mut w3 = fill;
|
||||
for (i, &x) in xs.iter().enumerate() {
|
||||
if i != 10 {
|
||||
w3 = scratch_rewrite(x, &w3);
|
||||
}
|
||||
}
|
||||
assert_ne!(w, w3);
|
||||
}
|
||||
Loading…
Reference in a new issue