igneum/docs/analysis/scratch-soundness.md

40 KiB

Layer 3 soundness: the per-warp scratch with read-modify-writes

5 October 2026 (night), cryptographer role, Counter ASIC 2.0 plan step 4 (docs/plans/counter-asic-2.md). Branch ca2-soundness on top of readwidth b970dda (the scratch as a class parameter, 32 or 128 KiB per warp). Tests: igneum-pow/tests/scratch.rs; Metal runs through proto-metal/packbench on the M5 Max; commands and counts in docs/bench-log.md (entry of the same date). Nothing here touches the lottery hash as shipped: variant 5 is behind LoadClass::scratch(k, kb) and is never emitted by generator version 2.

Every figure below is measured (machine, date, command named) or cited; "approximate" marks a figure from memory.

0. The five findings

# Question Finding Status
1 Is what is written uniform and beyond a chip's precomputation? The fill is a bijection of the lane nonce, the rewrite a bijection of the fold value in each word; written words show no bit bias over 3 to 12 million rewrites per class (worst 3.63 sigma of 6). The fill IS precomputable, by design, and at 64 slots 78.5 percent of reads are fill reads. sound as a function; see 2 for what that means
2 Does any short cut avoid the writes? No short cut inside a unit: a slot after d read-modify-writes needs all d fold values (replay test). But the live state is bounded by the read-modify-write count, not by the scratch size, because CPU verification resets the scratch per unit: 64 to 320 bytes per lane at scr2 to scr8, whatever the nominal 32 KiB, 128 KiB or 1 MiB. The named chip (cache mirror plus recompute) keeps that in SRAM at under 5 percent of its mirror and its gain does not move at any share under the 6 GB cap. NOT sound as an anti-chip layer
3 Is the verifier's one-warp simulation exact? Exact when the GPU's lazy per-unit tag is unique over the arena's life and the arena holds no stale tag. The kernels rely on this and neither host guarantees it (no clear at allocation, no clear at the 32-bit wrap of the tag counter, 16.4 minutes on a 5090). With the host contract of section 4.3 the simulation is exact: 14 edge packs twice, 200 fuzz packs, consecutive units on one warp and the wrap inside a launch all match the CPU on Metal (228 of 228); a broken tag and a broken fill are caught (3 of 3). sound with a host contract; today it is luck
4 The attack surface of the writes Out of bounds: impossible by the mask, 42 of 42 emitted kernels pass the static check, which catches six deliberate breaks. Aliasing: none, lane-major arenas disjoint by (warp, lane), two logical units of a wave64 get two arenas. Ordering: one lane, one slot, program order; no cross-lane sharing, no atomics needed. Alignment: 16-byte slots at 16-byte offsets from a 256-byte-aligned base. Wrap: identical to the CPU, tested at the launch level. sound
5 What a conformance vector must carry The class and geometry, the fill and rewrite, the host contract (tags, clearing, groups a multiple of warps), two consecutive units on one warp with a forced slot collision, a unit in the top 256 nonces with the wrap inside the launch, and the fingerprint declared independent of the warp count. The standard three-unit vectors catch a broken tag only through base 1,000,000 and would miss it at a 1 MiB scratch. defined in section 6

Recommendation (section 10): do not adopt layer 3 as the plan states it (read-modify-writes taken from the 16 dataset loads). It replaces latency-bound dataset reads with cache-bound ones for the GPU, costs the named chip nothing it cannot keep in a few megabytes of SRAM, and leaves that chip's gain at 2.4x at every share. The lever that moves that chip is the mixer multiplier of the M16 analysis (x2 brings it to 1.2x, x4 to 0.6x, under the verifier's 10 ms gate). If a scratch is kept for another reason, add the read-modify-writes beside the 128 loads, never in their place, and ship the host contract and the vector of section 6 with it.

1. What the branch implements

Piece Where What
Class igneum-pow/src/generator.rs:170-230 LoadClass { scratch: Some(k), scratch_kb }: k of the 16 memory slots are Op::Scratch; scratch_kb KiB per warp of 16-byte slots, lane-major, slots = kb x 2 per lane (32 KiB: 64, 128 KiB: 256); scratch_slot_mask() = slots - 1
Draw generator.rs:488-491 the first k of the 16 drawn load slots become scratch ops (a uniform k-subset); the source register follows the fresh-source rule like a load
Fill igneum-pow/src/verify.rs:30 scratch_fill(seed, base, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot x 0x9e3779b1 + (j + 1) x 0x85ebca77), j in 0..2
Fold verify.rs:18 x = dst ^ w0; x = (rotl(x, 11) x 0x9e3779b1) ^ w1; x = (rotl(x, 11) x 0x9e3779b1) ^ w2; dst = x (the read-width fold over the three data words)
Rewrite verify.rs:41 the slot becomes (x ^ w1, rotl(x, 7) ^ w2, x + w0)
CPU model verify.rs:48-100, :305-312 ScratchModel: per (lane, slot) a written bit and three words; an unwritten slot reads as its fill; one model per unit, so a unit starts from the fill
Acceptance igneum-pow/src/accept.rs:202-215, 374 a scratch site that reads one slot in all 32 lanes rejects the program (lane-constant site); scratch slots carry bit 31 in the address list and are left out of the distinct-address bound, which now covers the dataset loads only
GPU statement igneum-pow/src/emit.rs:143-157 s_ = rN & mask; v_ = 16-byte load of slot s_; m_ = (v_.x == tag) ? ~0 : 0; w = (v_.yzw & m_) | (fill & ~m_); fold; dst = x_; 16-byte store of (tag, x_ ^ w1_, rotl(x_, 7) ^ w2_, x_ + w0_) in Metal, CUDA and OpenCL
Persistent prologue emit.rs:159-175 lane = tid & 31; warp_ = tid >> 5; arena = scratch + (warp_ x 32 + lane) x words_per_lane; for (g_ = warp_; g_ < groups; g_ += nwarps_) { gbase = baseNonce + g_ x 32; tag = salt + g_; ... }
Hosts proto-metal/packbench.swift:144-164, proto-opencl/host.c:1025-1033, 1268 the arena is allocated and never written by the host; salt starts at 1 and advances by the launch's unit count; no clear at allocation, none at the wrap

The constraint of the night (coordinator, 5 October 2026): the whole working set on an 8 GB card stays under 6 GB (1 GiB table, the layer 5 hot table, the scratch of every resident warp, buffers), which caps the scratch at tens of KiB per warp. On an RTX 5090 at full occupancy (170 SMs x 64 warps = 10,880 warps, approximate hardware maximum; the measured version 2 kernel ran 24 warps per SM, 4,080 warps, docs/bench-log.md M11, 4 October 2026):

Scratch per warp 10,880 warps 4,080 warps (measured occupancy) Table + scratch at 10,880 Under 6 GB with a 1 GiB table
32 KiB 340 MiB 128 MiB 1,364 MiB yes
128 KiB 1,360 MiB 510 MiB 2,384 MiB yes
1 MiB (the first experiment) 10,880 MiB 4,080 MiB 11,904 MiB no

2. Question 1: uniformity of what is written

2.1 As functions

The fill of word j of slot s for lane nonce n is splitmix32(((n ^ seed[j]) + s x 0x9e3779b1 + (j + 1) x 0x85ebca77)). splitmix32 is a bijection of its 32-bit input; for fixed (seed, s, j) the input is a bijection of n. So over any 2^32 consecutive nonces every 32-bit value appears once as the fill of (s, j): uniform. Test fill_is_a_bijection_of_the_nonce: 2^16 consecutive nonces give 2^16 distinct words for 7 slots x 3 word positions; the fill of lane l at base b equals the fill of lane 0 at base b + l; it wraps with the nonce (base 0xffffffe0, lane 32 equals nonce 0).

The rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0) is, for fixed old content w, a bijection of the fold value x in EACH word. Test rewrite_is_a_bijection_of_the_fold_value: 2^16 consecutive x give 2^16 distinct words in each position for 16 random w. Consequence: a uniform x gives a uniform word in every position, and the three words are three images of the same x, so a rewritten slot carries exactly 32 bits of new state behind 96 bits of storage (from w and any one written word, x is recovered; the test checks all three inversions).

The fold value x is fold(dst, w), a bijection of dst for fixed w (xor, then rotate-multiply-xor twice; the multiplier is odd). So the written words are uniform whenever dst is, and dst is a register of the running program.

2.2 The attack: what a chip can precompute

The fill is a pure function of (seed, nonce, slot): precomputable, and meant to be (the verifier computes it too). A chip never stores a fill; it computes it in about 10 integer operations when a slot is first touched. The written words depend on dst, the register state at that instruction, which depends on every earlier instruction of the hash, including the dataset loads. Nothing about them is precomputable before the hash runs. This is the whole of what question 1 can give: the writes are as unpredictable as the registers. What that is worth is question 2.

2.3 The stats run (the TESTS.md section 3 shape)

Test written_words_unbiased_and_rehit_rates, M5 Max, 5 October 2026, cargo test --test scratch: for each class, programs of igneum-genesis, igneum-genesis/stats1, igneum-genesis/stats2, 2^11 units each (196,608 hashes per class), closed-form dataset, every read-modify-write traced (verify::interpret_warp_scratch). Ones count per bit of every written word and of the change each rewrite makes (written XOR read), sigma = sqrt(N)/2, limit 6 sigma like the acceptance rule's output check.

Class Slots per lane RMW per hash per lane Rewrites traced Max bias, written words (sigma) Max bias, written XOR read (sigma)
scr2k32 64 16 3,145,728 2.61 3.40
scr4k32 64 32 6,291,456 2.18 3.81
scr8k32 64 64 12,582,912 3.63 2.25
scr2k128 256 16 3,145,728 3.36 2.19
scr4k128 256 32 6,291,456 2.71 3.68
scr8k128 256 64 12,582,912 2.73 2.60

576 bit positions (6 classes x 3 words x 32 bits) at under 4 sigma is what fair coins give. Verdict: no structural bias in what is written. Like TESTS.md section 3 this is a sanity check, not a proof of strength.

3. Question 2: no short cut avoids the writes

3.1 Inside a unit: the chain is dependent

Slot s of lane l, touched d times in a unit, holds w_d = rewrite(x_d, w_{d-1}), w_0 = fill, with x_i = fold(dst_i, w_{i-1}). x_i depends on the slot content before it, which depends on every earlier fold value of that slot; and dst_i is the register state, which the earlier fold values entered. Test slot_is_replayable_from_its_fold_values: a slot after 64 read-modify-writes is reproduced from the fill and the 64 fold values; dropping one diverges. So a chip cannot skip a write and still read the slot later. It has three ways to hold a slot, all exact:

Store Bytes per lane Cost on a re-hit
Dense: every slot, 12 data bytes plus a valid bit 12 x slots: 776 (64 slots), 3,104 (256), 24,832 (2,048) one SRAM read
Sparse: only touched slots, 12 bytes plus a slot index about 13 x distinct: 185 to 820 (table below) one lookup
Implicit: only the fold values, 4 bytes plus a slot index per read-modify-write, replay on a re-hit 5 x 8k: 80 (scr2), 160 (scr4), 320 (scr8) d rewrites of 5 integer ops

The implicit store is smaller than the dense one whenever slots > 8k / 3: at scr4 above 10.7 slots, at scr8 above 21.3. So "the smallest scratch at which keeping it implicitly is dearer than storing it" is 8k/3 slots per lane, 2.7 to 5.3 KiB per warp at scr4 to scr8. Every size on the table, 32 KiB and above, is past it: a chip keeps the scratch implicitly in 80 to 320 bytes per lane at any nominal size, and the replay cost is bounded by the re-hit depth, which the next table measures.

3.2 The re-hit rate at 64 and 256 slots (and at 2,048)

Measured in the same test run (every read-modify-write of 196,608 hashes per class traced; a re-hit is a read of a slot the same unit wrote earlier). Birthday: distinct = S (1 - (1 - 1/S)^n) for n uniform draws from S slots.

Class S n = RMW per hash Distinct slots, birthday Re-hits, birthday Re-hit %, birthday Re-hit %, measured Max chain depth seen Slot histogram against uniform
scr2k32 64 16 14.26 1.74 10.9 12.58 7 chi2 z 22,023; hottest slot 2.74x, coldest 0.83x
scr4k32 64 32 25.33 6.67 20.8 21.47 8 z 10,880; 1.87x, 0.92x
scr8k32 64 64 40.64 23.36 36.5 36.99 9 z 7,587; 1.39x, 0.91x
scr2k128 256 16 15.54 0.46 2.9 3.84 5 z 19,146; 5.10x, 0.82x
scr4k128 256 32 30.14 1.86 5.8 6.25 6 z 9,632; 3.06x, 0.90x
scr8k128 256 64 56.72 7.28 11.4 11.89 6 z 5,873; 1.98x, 0.90x
1 MiB (not run) 2,048 32 31.76 0.24 0.8

Two readings. First, the slot a read-modify-write addresses is the low 6 or 8 bits of a program register, and those bits are not uniform: or sets them, mul clears them, so one slot of 256 is addressed 5.1 times as often as the mean and the re-hit rate runs 2 to 33 percent above the birthday rate. For the dataset the same bias on the low bits of a 28-bit address is harmless (it moves the read inside an item); for a 64-slot scratch it concentrates the chain. Second, the chain depth is small: at scr4k32 the deepest slot in 196,608 hashes saw 8 earlier read-modify-writes; a replay costs at most 8 x 5 integer operations, against about 1,170 for one dataset item.

3.3 The live state is bounded by the read-modify-write count, not by the size

The verifier evaluates one unit from nothing but (program, day, nonce group): ScratchModel::new per unit, verify.rs:296. Every conforming GPU must therefore start every unit from the fill, which the tag does (section 4). So no state crosses a unit boundary, and the state a unit can ever read back is what it wrote itself: at most 8k slots per lane. The nominal size only sets how often those 8k writes land on the same slot (the table above). The scratch's "memory" is 8k x 16 bytes per lane of touched slots, 256 bytes to 1 KiB at scr2 to scr8, and a chip holds it implicitly in 80 to 320 bytes.

The attack of rolling back or sharing scratch between units has nothing to take: a unit starts from the fill whatever ran before it, so a chip that clears 64 valid bits per unit has rolled back, and nothing one unit wrote is readable by another. The CPU verifier is that chip.

3.4 The named chip, and what the scratch costs it

The strongest chip the plan has priced (coordinator, 5 October 2026): the whole 256 MiB cache on the die, computing every dataset item on the fly. Its cache SRAM, from docs/analysis/sram-mirror.md revision 2 (ca2-analysis e6085c6), headline at shipped-product density / bit-cell lower bound, dollars per good die approximate: 164 / 83 mm^2 and $30 / $13 at N7 (shipped density from AMD 3D V-Cache, 64 MB on 41 mm^2, Hot Chips 2021); 128 / 64 mm^2 and $46 / $21 at N5, N3E and Intel 18A (TSMC N5 HD macro 31.8 Mib/mm^2 after assist overhead, SemiAnalysis, December 2022); 106 / 54 mm^2 and $56 / $26 at N2; with a 96 MB hot table 226 / 114 at N7, 175 / 89 at N5, 146 / 74 at N2. The chip's cache cost in the table below is the N5 headline, 128 mm^2 and $46 per good die. It computes every item through the mixer (docs/analysis/m16-recompute-attacker-2026-10-05.md: 128 items per hash, about 1,170 integer operations per item, 150,000 per hash; at a 50 T op/s integer budget equal to a 5090's, approximate, 0.33 Ghash/s). Against the measured version 2 rate of the RTX 5090, 139.7 MH/s (docs/bench-log.md M11, 4 October 2026), that is 2.4x before any fixed-function factor, 7x with the 3x the M16 analysis allows (approximate).

Units in flight on that chip. It has no DRAM latency to cover: every one of its 1,024 cache reads per hash is an on-die SRAM read. Its hash latency is the dependent chain: 128 items x (8 dependent SRAM reads plus 9 mixer applications). At about 10 ns per on-die read and about 40 ns per 130-operation mixer on a 16-wide integer pipeline at 2 GHz (both approximate), an item is about 0.4 us and a hash about 50 us; at 0.33 Ghash/s that is about 17,000 hashes in flight, 530 units of 32 lanes. A tighter pipeline halves it. The GPU covers DRAM latency (40 to 48 ns row cycle, MEMSYS 2018, more under load) with 130,560 lanes in flight at the measured occupancy (4,080 warps x 32), 348,160 at full occupancy, that is 8 to 20 times more lanes than the chip needs.

What the scratch costs that chip, per variant, with the arithmetic:

Chip cache mirror: 128 mm^2, $46 per good die (N5 headline; 64 mm^2, $21 bit-cell lower bound). Chip scratch SRAM at the same two densities (2.1 MB/mm^2 headline, 4.2 MB/mm^2 lower bound at N5):

Variant Dataset loads per hash Chip ops per hash Chip rate at 50 T op/s 5090 rate Chip gain Chip scratch SRAM at 17,000 lanes, implicit store Same, dense 64-slot store Dense store as mm^2, headline / lower bound (N5) Share of the 256 MiB mirror (any density)
scr0 (control), 128 loads 128 150,000 333 MH/s 139.7 measured 2.4x 0 0 0 0
12.5% replaced (scr2) 112 131,400 381 160 projected (128/112 x 139.7) 2.4x 1.4 MB 13 MB 6.2 / 3.1 mm^2 4.9%
25% replaced (scr4) 96 112,800 443 186 projected 2.4x 2.7 MB 13 MB 6.2 / 3.1 4.9%
50% replaced (scr8) 64 75,600 661 279 projected 2.4x 5.4 MB 13 MB 6.2 / 3.1 4.9%
12.5% added (16 RMW beside 128 loads) 128 150,200 333 139.7 or below 2.4x or more 1.4 MB 13 MB 6.2 / 3.1 4.9%
25% added 128 150,400 332 139.7 or below 2.4x or more 2.7 MB 13 MB 6.2 / 3.1 4.9%
50% added 128 150,800 332 139.7 or below 2.4x or more 5.4 MB 13 MB 6.2 / 3.1 4.9%
256-slot dense store (128 KiB class), any share 53 MB 25 / 12.6 20%

How the rows are computed: a read-modify-write costs the chip about 12 integer operations (fold and rewrite) and one SRAM access; replacing a load removes an item derivation (1,170 operations); the 5090's rate for a replaced load is projected from the measured distinct-load bound (the card's rate tracks distinct dataset loads per hash, docs/bench-log.md 3 October, 23.7 G loads/s at 1 GiB; the readwidth agent's M5 Max measurement of the night, relayed by the coordinator, shows the same: 27.7 MH/s at v2 to 29.4-31.7 at 25 percent replaced and 44.4-49.1 at 50 percent, 32 KiB per warp). The scratch SRAM is 17,000 lanes x 80 to 320 bytes (implicit) or x 776 bytes (dense at 64 slots) or x 3,104 bytes (dense at 256 slots); its share of the mirror is a ratio of bytes, 4.9 or 20 percent, whichever density is used for both; the implicit store (the chip's cheaper choice at every size, section 3.1) is 0.5 to 2 percent. The chip's gain is set by operations per dataset item and the GPU's distinct-load bound, and the scratch touches neither.

Plain answer to the coordinator's question: no read-modify-write share under the 6 GB cap, replaced or added, brings the named chip under 2x. The share would be chosen as the smallest at which the chip falls under 1.5x, and there is none: the gain is 2.4x at 0, 12.5, 25 and 50 percent, 32 or 128 KiB. This changes nothing about the public claim that layer 3 would have changed: the claim must rest on the mixer, not on the scratch.

The lever that does move that chip, from the M16 table, beside it:

Mixer cost multiplier Chip ops per hash Chip rate Gain against 139.7 MH/s, no fixed-function factor With a 3x factor (approximate) CPU verify per warp (M16 table, scaled from 0.41 to 1.2 ms) 5090 daily dataset build
x1 (today) 150,000 333 MH/s 2.4x 7.2x 0.4 to 1.2 ms 13.4 ms
x2 300,000 167 1.2x 3.6x 0.8 to 2.4 ms 27 ms
x4 600,000 83 0.6x 1.8x 1.6 to 4.8 ms 54 ms
x8 1,200,000 42 0.3x 0.9x 3.3 to 9.6 ms 107 ms

The mixer multiplier leaves the honest hash rate untouched (the miner pays the mixer once a day), costs the chip linearly, and is bounded by the 10 ms verification gate (x8 is at the gate's edge on this core, and the 2019-class core of O-1.14 is unmeasured). The scratch costs the honest GPU a measured share of its rate when it spills the cache and nothing when it does not, and costs the chip a few megabytes. The comparison is not close.

3.5 Where the GPU's writes would cost DRAM latency, and why that does not help

The GPU's hot scratch footprint is not the nominal size either: it is the slots in-flight units have touched, about warps x 32 lanes x distinct slots x 16 bytes (x 2 at a 32-byte sector, approximate): on the 5090 at 4,080 resident warps and scr4, 25.3 slots at 64 or 30.1 at 256, 53 to 63 MB of slots, 100 to 125 MB in sectors, around the card's 96 MiB L2 (docs/bench-log.md, 3 October). The readwidth agent's M5 Max rows (coordinator's message: the rate rises with the share at 32 and 128 KiB) show the scratch sitting in that chip's caches at 4,096 warps. To push the writes to DRAM latency the hot footprint must pass the last-level cache at the resident count: 96 MiB / 4,080 warps = 24 KiB per warp, which at 512 bytes of touched slots per lane per read-modify-write slot means 8k x 512 B > 24 KiB, k above 6 (above 48 read-modify-writes per hash) at ANY nominal size on the table, or a higher resident count. That fits the 6 GB cap (it is the hot set, not the arena, that matters), and it costs the honest miner a DRAM-latency read-modify-write per slot (a DRAM row cycle is 40 to 48 ns across DDR4, GDDR5 and HBM2, Li, Reddy and Jacob, MEMSYS 2018; the loaded latency a GPU kernel sees is higher, approximate; DRAM latency improved 1.3x in two decades while bandwidth improved 20x, Chang 2017, so no memory technology an attacker could buy removes it, and no shipped mining chip has used HBM or stacked memory) while the named chip still keeps the same hot set in a few megabytes of SRAM at 8 to 20 times fewer lanes in flight. The write path cannot be made to cost the chip more than the GPU, because the GPU must keep 8 to 20 times more of it live.

4. Question 3: the verifier's one-warp simulation is exact

4.1 Lazy fill on both sides

The CPU initialises lazily with a written bit per (lane, slot), one model per unit. The GPU initialises lazily with a 32-bit tag in word 0 of each 16-byte slot: a slot whose tag equals the unit's tag reads as written, any other reads as the fill (emit.rs:143-157). There is no explicit fill and no reset between units of a persistent warp (emit.rs:159-175: the loop over g_ keeps the arena). The two agree if and only if, when a unit first touches a slot, that slot does not already carry the unit's tag. That is:

  1. Tags are unique over the life of the arena's contents (tag = salt + g_, salt the host's running counter).
  2. The arena holds no word equal to a live tag in a slot's tag position before the unit writes it.

4.2 The attacks (the bug classes)

Case What happens Today
Recycled allocation A fresh process starts salt at 1 (packbench.swift:144, host.c:1027). If the driver hands back the previous process's arena with its contents (Metal, CUDA and OpenCL do not promise zeroed memory, approximate), slots tagged 1..N from the old run match the new run's first units exactly, and those units read stale words instead of the fill: a CPU mismatch on every colliding slot. not guarded; passes on this Mac because fresh allocations read as zero in practice and tag 0 is never issued (luck, not contract)
Tag counter wrap salt is 32 bits and advances by units per launch. A 5090 at 139.7 MH/s runs 4.37 M units/s, 2^32 units in 984 s: the counter wraps every 16.4 minutes on one card (81.8 minutes on the M5 Max at 28 MH/s). After the wrap a slot whose LAST writer carried the repeated tag reads as written. With 10,880 arenas each slot is rewritten about 395,000 times between two uses of one tag (at 64 slots a unit leaves a slot untouched with probability 0.60; 0.60^395,000 is 0), so on a full card the wrap is harmless in practice; on a one-warp launch repeated 2^32 times it is not. not guarded
Tag 0 on zeroed memory A host that starts salt at 0 gives unit 0 the tag 0, which a zeroed arena carries in every slot: unit 0 reads zeros for every first touch. both hosts start at 1; nothing in the pack says they must
groups not a multiple of the warp count Warps run different trip counts; the OpenCL local-memory exchange path carries a barrier inside the loop (spec 1.9), so a short warp hangs or desynchronises. packbench refuses it; host.c rounds the batch

4.3 The host contract that makes the simulation exact

A host of a scratch class MUST: allocate the arena as warps x 32 x words_per_lane words and zero it; issue tags from a 32-bit counter that starts at 1 and advances by the unit count of every launch; zero the arena again before any launch whose tags would pass 2^32 - 1 (tag 0 is never issued); launch groups as a multiple of the warp count. The zeroing costs one memset of the arena (340 MiB at 32 KiB x 10,880 warps) every 2^32 units, 16 minutes on a 5090. This is the class fix for all four rows: with it the GPU's tag test and the CPU's written bit are the same predicate.

4.4 The tests (Metal, M5 Max, 5 October 2026)

Two consecutive units on one persistent warp and the wrap inside a launch (packbench --warps 1, --batch-base 4294967040, the option added on this branch); the hand-built edge programs that force every read-modify-write of a hash onto one slot (so two consecutive units on one arena collide on every slot); the deliberate breaks. Results in section 7.2. On the CPU, the same edge programs against an independent hand model (a second interpreter with its own slot store, tests/scratch.rs): 56 of 56 cases match, and the hand model with its rewrite words swapped mismatches on every case (the comparison has teeth).

5. Question 4: the attack surface of the writes

Surface Argument Test
Out of bounds s_ = rN & (slots - 1), so s_ < slots; the lane's arena is (warp_ x 32 + lane) x 4 x slots words from the base, the access is arena + 4 x s_ + 0..3, the largest index is warps x 32 x 4 x slots - 1, the host's allocation. The emitter has one scratch template (emit.rs:143) and it masks. scr_packs_regenerate_and_pass_the_static_scratch_check: 42 of 42 emitted kernels (7 scr packs x 6 files, the OpenCL bound file carrying two kernels) regenerate byte for byte from program.json and pass the text check: k masked slot definitions with the class mask, k tagged stores, 3k fill calls, one arena definition with the class stride, one tag definition, no scratch[; six deliberate breaks caught (section 8)
Aliasing between lanes Lane-major: lane l of warp w owns words [(32w + l) x 4S, (32w + l + 1) x 4S); two (w, l) pairs give disjoint ranges. Inside the range a slot is 4 words at 4 x s_, so two slots of one lane are disjoint too. the lanevar edge program: one init-dependent slot per lane, 32 lanes at 64 slots share slots in pairs by the birthday bound; any cross-lane aliasing would change the fold; 128 of 128 lanes on Metal (section 7.2)
Wave64 (two logical units in one hardware wave) warp_ = tid >> 5, so the two halves get warp_ = 2w and 2w + 1, two arenas; gbase and tag are per g_, per half. not run on wave64 hardware (the OpenCL emulator's persistent launch is on the readwidth commit; unverified here)
Determinism: alignment A slot is 16 bytes at byte offset 16 x (lane_base + s_); the arena base is the buffer base: Metal, CUDA and OpenCL allocations are at least 128-byte aligned (CUDA 256, OpenCL CL_DEVICE_MEM_BASE_ADDR_ALIGN at least the largest built-in type, approximate from memory), so every 16-byte vector access is aligned. Metal: every run of section 7
Determinism: ordering A lane's two read-modify-writes of the same slot in one hash are a load and a store, then a load and a store, from one thread to one address: program order within a thread holds in every model. No other thread touches the slot (aliasing row), so no atomics, fences or barriers are needed and none are emitted. slot0 and sixteen edge programs: 64 and 128 dependent read-modify-writes on one slot per lane per hash, standalone and as the second unit on a warp
Determinism: vendors The statement is integer only: xor, rotate by immediate, multiply, add, a 16-byte load and store. Bit-exact across Metal, CUDA and OpenCL by construction; measured only on Metal here. Metal; CUDA and OpenCL runs are PC jobs (not mine tonight)
32-bit nonce wrap gbase = baseNonce + g_ x 32 and nonce = baseNonce + gid wrap in 32-bit arithmetic; scr_fill(gbase + lane) wraps like the CPU's base.wrapping_add(lane); out[gid] indexes by launch position, not by nonce. An aligned unit never straddles 2^32 (spec 1.9), so the wrap case is a launch whose unit SEQUENCE crosses it. packbench --batch-base 4294967040 --batch-log2 9: 16 units from 0xffffff00, the ninth at gbase 0; fingerprint identical at 1 and 4 warps (section 7.2); every fuzz pack runs that launch

6. Question 5: what a vector for the scratch class must carry

Before a scratch pack can be a conformance vector (plan step 4, "only then a vector"), it must carry, beyond what igneum-program-pack-3 carries today:

  1. The class in the program id and the pack (scr<k>k<kb>: it is, program_id_class, generator.rs:400-412) and the geometry (slots per lane, words per lane, bytes per warp: it is, program.h).
  2. The fill and the rewrite as text (it is, program.json "scratch").
  3. The host contract of section 4.3 as text in program.h and program.json: tag counter from 1, zero at allocation and at the wrap, groups a multiple of the warp count. Not there today.
  4. Vectors that exercise the tag path, which the three standard units do not reliably: two consecutive units on one warp (bases 0 and 32 in one one-warp launch) for a program whose consecutive units collide on a slot. At 64 slots any generated program collides (25 touched of 64 per unit; the broken-tag run of section 8 was caught by base 1,000,000, a warp's 16th unit, and NOT by a two-unit launch whose vectors lack base 32). At 2,048 slots two consecutive units share a touched slot with probability about 0.4 (32 x 32 / 2,048 expected overlaps = 0.5), so the standard vectors would miss a broken tag at the 1 MiB size with probability about 0.6 per unit pair. The edge programs slot0 and sixteen collide on every slot at every size: a vector set should carry one.
  5. A unit in the top 256 nonces with the launch crossing 2^32 (--batch-base near the top, at least two warps).
  6. The batch fingerprint declared independent of the warp count (8c07620f4d9adefd for scr4k32 at 2^12 nonces from base 0 at 1, 2 and 128 warps, section 7.2): unit independence is the property the per-unit reset gives, and a fingerprint that moved with the warp count would mean a unit read another unit's slot.

7. Tests and results

7.1 CPU (igneum-pow/tests/scratch.rs, cargo test -j4 --test scratch, M5 Max, 5 October 2026, 3.6 s)

Test What Result
rewrite_is_a_bijection_of_the_fold_value 16 random slot contents x 2^16 consecutive fold values, each written word distinct; the three inversions pass
fill_is_a_bijection_of_the_nonce 7 slots x 3 words x 2^16 nonces distinct; lane and base interchange; wrap pass
written_words_unbiased_and_rehit_rates 6 classes x 3 seeds x 2^11 units, every rewrite traced: bias within 6 sigma (worst 3.63), re-hit rate within 0.9x to 2x of birthday, slot histogram, depth histogram pass (tables of sections 2.3 and 3.2)
edge_programs_match_the_hand_model 7 edge programs x 2 geometries x 4 bases (0, 32, 0x7ffffff0, 0xffffffe0) against an independent hand model; the slots driven and the re-hit counts as built; the mutated hand model mismatches 56 of 56 pass, 56 of 56 teeth
scr_packs_regenerate_and_pass_the_static_scratch_check 7 scr packs: program and program id from program.json, 6 kernel texts byte for byte, static scratch check on all 42, the pack's vectors from the CPU; six deliberate breaks caught pass
fuzz_scr_programs_cpu 200 generated programs over the six classes, generator contract and acceptance on every one, 4 units each (one in 0..224, one around 2^31, one in the top 256 nonces, one uniform), traced run equal to the untraced run, every slot inside the lane; writes the 214 packs for Metal with IGNEUM_SCRATCH_PACKS_OUT pass; 200 of 200 have a unit in the top 256
slot_is_replayable_from_its_fold_values 64 dependent read-modify-writes replayed from the fill and the fold values; one dropped diverges pass

The rest of the crate: 33 of 34 lib tests and all pack tests pass; verify::tests::fold_and_wide_fetch fails on the readwidth tip itself (verify.rs:508, k as u32 * 0x9E37_79B1 overflows under the test profile's overflow checks; the readwidth agent's test, reported to its owner, not touched here).

7.2 Metal (proto-metal/packbench built from this branch, M5 Max, 5 October 2026, under with-lock.sh run)

Run Launch Expected Result
scr4k32, standard pack 2,048 warps, 2^24 nonces, 1 batch 3 of 3 standalone, 3 of 3 in batch PASS, fingerprint 3d1af881bd978fb9; 1.8 s wall for the whole run (compile, cache, 1 GiB build, vectors, batch)
scr4k32, warp-count independence 2^12 nonces (128 units) at 1, 2 and 128 warps one fingerprint 8c07620f4d9adefd at all three, PASS
scr4k32, wrap inside the launch 512 nonces from 0xffffff00 at 1 and 4 warps one fingerprint, the base-0 vector inside the window after the wrap 8e9e233234d3a297 at both, in-batch 1 of 1, PASS
scr4k32, broken tag (tag = salt), standard vectors 2,048 warps, 2^24 the base-1,000,000 vector (warp 530's 16th unit) fails standalone 3 of 3, in batch 2 of 3, overall FAIL (caught)
scr4k32, broken tag, two units on one warp 1 warp, 2^6 nothing to catch it: the standard vectors have no base 32 standalone 3 of 3, in batch 1 of 1, PASS (missed: the point of section 6 item 4)
14 edge packs (7 programs x 32 and 128 KiB), run A 1 warp, 2^6 (units at bases 0 and 32 on one arena) 4 of 4 standalone, 2 of 2 in batch each 14 of 14 PASS (56 of 56 standalone units, 28 of 28 in batch)
14 edge packs, run B 1 warp, 2^9 from 0xffffff00 (16 units on one arena, the wrap inside) 4 of 4 standalone, 3 of 3 in batch each 14 of 14 PASS (56 of 56, 42 of 42)
edge slot0 at 32 and 128 KiB, broken tag (tag = salt) 1 warp, 2^6 standalone 4 of 4, in batch 1 of 2, FAIL as expected at both geometries: the second unit read the first's slot 0 and FAILED; the standalone units passed
edge slot0 at 32 KiB, broken lazy fill (m_ forced to all ones: a first touch reads the stale words) 1 warp, 2^6 standalone fails 0 of 4 standalone, 0 of 2 in batch, FAIL (lane 0 of base 0: GPU 64b49aeb987dae69, expected 9ff3a2021f66b5be)
200 fuzz packs (scr2k32 29, scr4k32 26, scr8k32 42, scr2k128 32, scr4k128 26, scr8k128 45; datasets 64 MiB, 256 MiB, 1 GiB) 2 warps, 2^9 from 0xffffff00 (8 units per warp, the wrap inside) 4 of 4 standalone, 2 of 2 in the window, 200 of 200 PASS 200 of 200 PASS: 800 of 800 standalone units (25,600 hashes), 400 of 400 in batch; 91 s for the 200 runs

Totals on Metal: 228 of 228 runs PASS where a pass was expected, 3 of 3 FAIL where a failure was built in.

8. Deliberate breaks (the watcher rule)

Break Where Caught by Evidence
One slot mask dropped (Metal) copy of scr4k32 program.metal static check: "masked slot followed by the load: 3, expected 4" test output
Mask 63 changed to 127 on every RMW (Metal) same "masked slot followed by the load: 0, expected 4" test output
Arena stride 256 changed to 128 words (Metal) same "arena definition: 0, expected 1" test output
A stray arena[0] and scratch[1] access (Metal) same "arena mentions: 10, expected 9; direct scratch indexing: 1, expected 0" test output
One slot mask dropped (OpenCL, CUDA) copies of scr4k32 kernel.cl, kernel.cu "masked slot followed by the load: 3, expected 4" test output
Wrong class geometry or RMW count or kernel count passed against a right text the same text the check fails test output
tag = salt (every unit of a launch shares the tag) copy of scr4k32 program.metal, on the GPU the base-1,000,000 vector in a 2,048-warp batch vectors standalone 3/3, in batch 2/3, overall FAIL
the same on the slot0 edge pack, two units on one warp on the GPU in-batch 1 of 2 bench-log entry
lazy fill broken (m_ all ones) copy of the slot0 edge pack, on the GPU standalone vectors bench-log entry
The hand model's rewrite words swapped tests/scratch.rs every edge case mismatches 56 of 56

The out-of-bounds break (mask dropped) was not run on the GPU on purpose: Metal does not bounds-check device buffers (TESTS.md section 5), so a run would read another lane's or another buffer's words and "did not crash" would prove nothing. The static check is the guard, as it is for the dataset mask.

9. What is unverified

  1. CUDA and OpenCL runs of the scratch packs on NVIDIA and AMD (PC jobs, reserved for the readwidth agent tonight); the 5090's rate per variant, so the "projected" column of section 3.4 is the distinct-load bound, not a measurement. Wave64 hardware for the two-arena argument.
  2. The chip-side latency figures of section 3.4 (10 ns SRAM read, 40 ns mixer) are approximate; the conclusion does not depend on them: at ten times the in-flight count the scratch is still under a sixth of the mirror.
  3. The recycled-allocation case was not reproduced (it needs a driver that hands back live contents); the argument is that nothing forbids it and the contract of 4.3 removes it.
  4. The slot-bias finding (section 3.2) was measured on three seeds per class; the hottest-slot ratio will vary by program.
  5. verify::tests::fold_and_wide_fetch on the readwidth tip (section 7.1).

10. Recommendation

  1. Layer 3 is sound as a construct: the written words are uniform, the chain inside a unit has no short cut, the kernels cannot write out of bounds, and with the host contract of section 4.3 the CPU's one-warp simulation is exact (14 edge packs, 200 fuzz packs, the wrap, consecutive units on one arena, on Metal).
  2. Layer 3 is not sound as a chip-resistance layer, at the capped size or at any size: CPU verification resets the scratch per unit, so its live state is 8k slots per lane whatever the arena, a chip keeps it implicitly in 80 to 320 bytes per lane, and the named chip (on-die cache mirror plus recompute) keeps its whole scratch in 1.4 to 13 MB of SRAM at 530 units in flight, 3 to 5 percent of its mirror. Its gain stays at 2.4x (7x with a 3x fixed-function factor, approximate) at 0, 12.5, 25 and 50 percent, replaced or added, 32 or 128 KiB. No share under the 6 GB cap brings it under 2x.
  3. Taking the read-modify-writes from the 16 dataset loads makes the hash less memory-hard for everyone: the GPU measured faster at every share on the M5 Max (readwidth rows), and the chip's operations per hash fall with the loads. If a scratch is kept at all, add it beside the 128 loads. There is no reason found here to keep one.
  4. The lever that moves the named chip is the M16 mixer multiplier: x2 to 1.2x, x4 to 0.6x against the measured 5090 rate, at 0.8 to 4.8 ms of verification per warp against the 10 ms gate. Decision 2 should price that against the gate on the 2019-class core (O-1.14) rather than layer 3.
  5. If the project lead keeps layer 3 for a reason outside this analysis: ship the host contract in the pack, add the four vector items of section 6 (consecutive units with a forced collision, the wrap launch, the warp-count-independent fingerprint, the contract text), and run the CUDA and OpenCL twins of section 7.2 on the PCs before the class becomes a genesis rule.