LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.
Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
266 lines
33 KiB
Markdown
266 lines
33 KiB
Markdown
# Hot table: a second table sized to GPU cache, read beside the 1 GiB dataset
|
|
|
|
Counter ASIC 2.0, layer 5 (`docs/plans/counter-asic-2.md`). Experiment branch `ca2-cache`, 5 October 2026 (night), on top of the read-width branch (`readwidth` b970dda: `LoadClass`, `verify::fold_words`, the scratch variant, `proto-metal/packbench.swift`, `proto-opencl/host.c --bench-pack`). Nothing here is the lottery hash: every hot class sits behind the generator flag and the default v2 path is byte-identical (the pinned packs `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` are diffed by `igneum-pow/tests/packs.rs`).
|
|
|
|
## 1. The idea
|
|
|
|
The honest hash is 128 dependent random 4-byte reads over 1 GiB (spec 01 section 1.5). A chip that wants a gain on that must beat a GPU at DRAM random reads, which is the same physics for both (the plan's last section). What a chip can do that a GPU cannot is choose its memory: a chip can mirror read-only data into SRAM and serve it at SRAM latency, if the data fits.
|
|
|
|
The hot table turns that around. A second table `H` of `S` MiB (32, 64 or 96 in this experiment) is derived from the epoch seed and read by `k` of the 16 load slots (2, 4 or 8). `S` is chosen to fit the caches of the cards that mine: 96 MiB L2 on the RTX 5090, 64 MB Infinity Cache plus 8 MB L2 on the RX 9070 XT, the system level cache on the M5 Max (figures approximate, from memory; the probe of section 6 measures what each card does at each size). A GPU gets the hot loads as cache hits for free. A chip must carry `S` MiB of SRAM for the same hits, beside the DRAM path it still needs for the other `16 - k` slots, or serve `H` from DRAM and fall behind the GPU by the hot share. Either way the chip pays for something the GPU already has.
|
|
|
|
What this does not change: the cold loads (the 1 GiB dataset) stay a dependent chain at DRAM latency, so the latency-bound property holds for them; the hot loads are interleaved in the same chain (every load's address is a fresh register of the same iteration, G2), so a hot hit shortens the chain by one DRAM latency and nothing else.
|
|
|
|
## 2. The specification text (proposed; prototype values, to be fixed at gate 1)
|
|
|
|
### 2.1 Hot key and fill
|
|
|
|
For an epoch whose program seed bytes are `e` (spec 01 section 1.12: the UTF-8 of a seed string in the packs, the 32-byte VDF output on the chain; the bytes before the attempt suffix of 1.4.6, so every attempt of one epoch shares one table):
|
|
|
|
```
|
|
KH = seed_words_from_bytes("igneum-hot/" || e) 8 words
|
|
```
|
|
|
|
`H` has `N_H = S x 2^18` words (`S` MiB) in `N_seg = S x 256` segments of 64 chained lines of 16 words, filled exactly as the cache of section 1.8.3 with `KH` in place of `K` and the tag `("Igne", "umHT") = (0x49676e65, 0x756d4854)` in place of `("Igne", "umMH")`:
|
|
|
|
```
|
|
prev = 0^16
|
|
for j in 0..63:
|
|
in = prev XOR (sigma[0..3] || KH[0..7] || s || j || tag[0..1])
|
|
line = B(in) ChaCha12 with feed-forward, section 1.8.2
|
|
H[segment s, line j] = line
|
|
prev = line
|
|
```
|
|
|
|
One GPU thread per segment, as the cache fill. The fill is a once-per-epoch cost: `S / 256` of the 256 MiB cache fill (section 1.8.3 table: 0.6 to 2.1 ms on the M5 Max GPU, 0.67 ms on the RTX 5090, 175 to 190 ms on one CPU core for 256 MiB), so under 1 ms on a GPU and 22 to 71 ms on one core, measured below.
|
|
|
|
### 2.2 The hot load
|
|
|
|
A hot load slot reads `H` instead of the dataset with the same fold (width 1: a plain XOR):
|
|
|
|
```
|
|
hot: dst = dst XOR H[mulhi(src, N_H)]
|
|
```
|
|
|
|
`mulhi(a, b)` is the high 32 bits of the 64-bit product (the `mulhi` family of section 1.4.1, bit-exact on Metal, CUDA and OpenCL). The index lies in `[0, N_H)` for any `N_H`, which is what lets `S = 96` exist: 96 MiB is not a power of two, so `src AND MASK` cannot address it. This is the multiply-shift range reduction section 1.13.3 proposes for the growing dataset, so the hot table is also its first measured use. For `S = 32` and `64` the mapping takes the top 23 or 24 bits of `src` where the dataset load takes the low 28; a fresh source is uniform, so neither choice costs uniformity (section 6, hot-load uniformity test).
|
|
|
|
The emitted text has one form per dialect, checkable by text search as the mask check of section 1.14 item 2: `hot[mulhi(rN, HOT_WORDS)]` (Metal), `hot[__umulhi(rN, HOT_WORDS)]` (CUDA), `hot[mul_hi(rN, HOT_WORDS)]` (OpenCL), with `HOT_WORDS` a literal of the pack. A hot pack's hash kernels carry exactly `16 - k` masked dataset loads and exactly `k` hot loads (`igneum-pow/tests/packs.rs`).
|
|
|
|
### 2.3 Which slots are hot
|
|
|
|
The 16 load slots are drawn first by the partial Fisher-Yates of section 1.4.3, which emits them in a uniformly random order. The first `k` slots in that draw order are the hot slots. Drawn, not fixed, because a fixed pattern (every fourth load, say) would let a pipeline schedule its SRAM reads statically for every hour; drawn costs no extra draw, so a `hot(S, k)` program takes the version 2 stream exactly and is the version 2 program with `k` of its loads redirected (the width roll of the read-width classes is not taken: `LoadClass::takes_width_roll`). A hot slot keeps every rule of a load: the fresh-source draw (G2), the injecting-write count (acceptance (b)), the lane-constant test and the distinct-address count (acceptance (c)); its address is tagged apart from dataset addresses in the count so a hot word and a dataset word at the same index are two addresses.
|
|
|
|
The class composes with the read-width fields of `LoadClass` (width mix, scratch `k` and `kb`): the hot slots are taken from the drawn slots after the scratch slots, so a class may carry a width mix, a scratch and a hot table at once. The experiment packs below use the version 2 widths and no scratch.
|
|
|
|
Two forms (coordinator's decision, 5 October 2026, after the first Mac measurement):
|
|
|
|
| Form | Load slots drawn | Dataset loads per hash | Hot loads per hash | Items per warp (verifier bound) | Class name | What a chip pays |
|
|
|---|---|---|---|---|---|---|
|
|
| replaced (`hot(S, k)`) | 16, the first `k` in draw order hot | `(16 - k) x 8` | `8k` | `(16 - k) x 256` | `hot64k4` | the on-die-cache recompute chip of `docs/analysis/scratch-soundness.md` 3.4 (the 256 MiB cache in SRAM, items derived on the fly, 333 MH/s at 50 T op/s, 2.4x the 5090) GAINS: a hot load replaces an item derivation (1,170 ops) with an SRAM read, so at `k = 4` its rate rises 1.33x against the GPU's measured 1.05 to 1.22x |
|
|
| added (`hot(S, k, added)`) | `16 + k`, the first `k` in draw order hot | `16 x 8 = 128` | `8k` | 4,096, unchanged | `hot64k4a` | `S` MiB of SRAM and `k` reads per iteration for nothing: the 16 item derivations stay; the GPU pays `k` cache hits |
|
|
|
|
The added form is the one that taxes the named chip; the replaced form stays as the measured record (section 6). In the added form the slot draw is a partial Fisher-Yates of `16 + k` slots over 1..63, so the program stream differs from version 2 (another slot count), and the width roll is still not taken (`takes_width_roll`: the slot count is 16 plus the class's own added hot slots). `loads_per_hash` is `128 + 8k`; the acceptance rule's distinct-address bound scales with it as for the read-width classes.
|
|
|
|
Program id: `FNV-1a 64 over "igneum-program-rw/" || generator || seed words || attempt || mix || load_slots || "hot/" || S || k [|| "added"]` (the read-width id with the hot fields appended), so no hot pack can be mistaken for a version 2 one, for another hot class or for the other form.
|
|
|
|
### 2.4 Acceptance (section 1.4.6)
|
|
|
|
The dynamic test stays a pure function of the program. A hot load reads `dataset_elem(mulhi(src, N_H), SW[2], SW[3])` (the six-operation closed form keyed by seed words 2 and 3, where the dataset stand-in is keyed by words 0 and 1), so the two stand-ins are two tables without any cache or day. Every other test is unchanged; the distinct-address bound is the read-width rule's (dataset and hot loads counted, scratch read-modify-writes not).
|
|
|
|
### 2.5 The verifier (section 1.11)
|
|
|
|
A verifier holds, per day, the mixer parameters and the 256 MiB cache, and, per epoch, `H` (`S` MiB, filled on one core in the times of section 6). A hot load is one table read per lane; a dataset load is the item derivation of section 1.11 as before. The verifier never holds the dataset. Per epoch the verifier's memory is `256 MiB + S MiB`.
|
|
|
|
## 3. Memory budget (the project lead's cap, 5 October 2026: the working set on a card stays under 6 GB on an 8 GB card)
|
|
|
|
Decided by the coordinator on 5 October 2026 after the card-lifetime review (`docs/analysis/card-lifetime-2026-10-05.md`, branch `card-lifetime` 1fecfe2, merged into `ca2-coord`): the GPU frees the 256 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms of fill and 13.4 ms of build on the 5090, bench-log 3 October 2026), so the cache is not resident and the working set is dataset + hot table + scratch (0 in v3) + buffers. The table below is that review's per-tier table with the hot table at its largest (96 MiB), buffers 128 MiB (the harnesses' 2^24-nonce output), resident warps = SMs x 48 on NVIDIA Ampere and later (SM counts approximate, from memory; the GTX 1650 is Turing at 32 warps per SM), 2,048 launched warps on Apple (the Metal harness). The scratch columns are the read-width variant's 32 and 128 KiB per resident warp, kept for the record; v3 carries no scratch.
|
|
|
|
| Tier | Card assumed (SMs, approximate) | Resident warps | Scratch at 32 KiB | Scratch at 128 KiB | Hot table | Dataset at genesis | Buffers | Total, no scratch | Total at 32 KiB | Total at 128 KiB | Years the dataset leaves under mapping (b), cache freed |
|
|
|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
| 4 GB | GTX 1650 (14 SMs x 32) | 448 | 14 MiB | 56 MiB | 96 MiB | 2,048 MiB | 128 MiB | 2,272 MiB | 2,286 MiB | 2,328 MiB | to year 4 (the 4 GiB step) |
|
|
| 8 GB | RTX 3050 (20) | 960 | 30 | 120 | 96 | 2,048 | 128 | 2,272 | 2,302 | 2,392 | to year 12 (the 8 GiB step) |
|
|
| 12 GB | RTX 3060 (28) | 1,344 | 42 | 168 | 96 | 2,048 | 128 | 2,272 | 2,314 | 2,440 | to year 28 (the 16 GiB step; year 12 if the cache were resident) |
|
|
| 16 GB | RTX 5060 Ti (36) | 1,728 | 54 | 216 | 96 | 2,048 | 128 | 2,272 | 2,326 | 2,488 | to year 28 |
|
|
| 24 GB | RTX 4090 (128) | 6,144 | 192 | 768 | 96 | 2,048 | 128 | 2,272 | 2,464 | 3,040 | to year 60 (the 32 GiB step) |
|
|
| 32 GB | RTX 5090 (170) | 8,160 | 255 | 1,020 | 96 | 2,048 | 128 | 2,272 | 2,527 | 3,292 | to year 60 |
|
|
| Apple 8 to 64 GB | M-series, 2,048 launched | 2,048 | 64 | 256 | 96 | 2,048 | 128 | 2,272 | 2,336 | 2,528 | 8 GB to year 4, 16 GB to year 12, 32 GB to year 28, 64 GB to year 60 (50% of unified memory usable, the review's assumption) |
|
|
|
|
The prototype packs here use the 1 GiB dataset of spec 1.5 (1,024 MiB less in every total). Every total is under 6 GB with the hot table at its largest, so `S` is not what the cap binds: the dataset's growth is, and the hot table takes 96 MiB of the room at every tier (about 2 months of the 0.5 GiB-a-year schedule). The verifier's memory is `256 MiB + S MiB` (section 2.5); the GPU's is the table above.
|
|
|
|
## 4. The chip model with the SRAM it would need
|
|
|
|
Figures: SRAM area per bit from `docs/analysis/m16-recompute-attacker-2026-10-05.md` section 3 (256 MiB in about 100 to 300 mm^2 at a current node: the low end from a 0.02 um^2 bit cell with array overhead, the high end from wafer-scale parts at about 1 MB per mm^2; approximate, from memory) and from `docs/plans/counter-asic-2.md` (256 MB in about 45 mm^2 at a leading node, approximate). That is 0.18, 0.39 and 1.17 mm^2 per MiB. Die areas: a 750 mm^2 GPU-class die (the M16 model's equal-silicon comparison) and a 100 mm^2 memory-chip die (an assumption for a latency-bound chip whose die holds memory controllers and little else; labelled as such).
|
|
|
|
| S (MiB) | SRAM at 0.18 mm^2/MiB | at 0.39 | at 1.17 | Share of a 100 mm^2 die (0.39) | Share of a 750 mm^2 die (0.39) |
|
|
|---|---|---|---|---|---|
|
|
| 32 | 6 mm^2 | 12 mm^2 | 37 mm^2 | 11% | 1.6% |
|
|
| 64 | 12 mm^2 | 25 mm^2 | 75 mm^2 | 20% | 3.2% |
|
|
| 96 | 17 mm^2 | 37 mm^2 | 112 mm^2 | 27% | 4.8% |
|
|
|
|
The gain arithmetic. Let `G0` be a chip's gain over a GPU on the version 2 hash (the plan's public claim: under 2x; the plan's last section says the bound is DRAM latency, the same physics on both). Let `g` be the GPU's own measured speed-up from the hot class over version 2 on the same card (section 6: `hash rate hot(S, k) / hash rate v2`). The ideal `g` with every hot load a hit at zero cost is `16 / (16 - k)`: 1.14x at `k = 2`, 1.33x at `k = 4`, 2.0x at `k = 8`; the measured `g` says what share of that a real cache delivers while the 1 GiB dataset streams through the same cache.
|
|
|
|
| Chip | Hot loads served from | Gain after the hot table | Arithmetic |
|
|
|---|---|---|---|
|
|
| A, no SRAM for H | DRAM | `G0 / g` | the chip's rate is what it was; the honest GPU gained `g` |
|
|
| B, S MiB of SRAM for H | SRAM | `G0 x A_die / (A_die + A_S)` per unit of silicon | the chip regains `g` and pays `A_S` on top of its die |
|
|
| B on a 100 mm^2 die, S = 64, 0.39 mm^2/MiB | SRAM | `0.80 x G0` | 100 / 125 |
|
|
| B on a 750 mm^2 die, S = 96, 0.39 mm^2/MiB | SRAM | `0.95 x G0` | 750 / 787 |
|
|
|
|
What the big-table loads still cost the chip: `(16 - k) x 8` dependent DRAM reads per hash at the card's loaded latency (the 9070 XT entry's probe: 451 ns on the 5090 at 4,096 lanes, 276 ns unloaded on the 9070 XT, 1,949 ns on the M5 Max at 4,096 lanes; section 6 repeats the probe here). At `k = 4` that is 96 reads per hash; to match one RTX 5090 at its honest 229 Mhash/s (the M16 model's reference) a chip must keep `229 M x 96 x 300 ns = about 6,600` DRAM reads in flight at a 300 ns latency, and 11,000 at 500 ns, whatever its arithmetic. That queue depth is a memory-controller property, which is the latency-bound argument of the plan restated for the cold share.
|
|
|
|
What the hot table does to the recompute attacker of M16: nothing good. `H` is a plain ChaCha12 chain, so a word of it can be recomputed from `KH` at `j + 1` block evaluations (32.5 on average, about 40,000 integer operations per hot load, approximate, against 1,170 per dataset item), which is 35x the cost of recomputing a dataset word. A chip stores `H` or reads it from DRAM; it does not recompute it.
|
|
|
|
Reading, before the measurements: the hot table costs a chip `A_S` of die area or `g` of rate. Which one binds depends on the measured `g`, which is why the measurement comes first. If a GPU's cache delivers most of the ideal `g` with the dataset streaming beside it, `k = 8` at `S = 64` doubles the honest rate and halves chip A. If the cache delivers little, the hot table is a cost to the verifier (`S` MiB per epoch) with no gain, and the layer is dropped.
|
|
|
|
## 5. Implementation (branch `ca2-cache`)
|
|
|
|
| Where | What |
|
|
|---|---|
|
|
| `igneum-pow/src/generator.rs` | `LoadClass { hot: Option<HotClass> }` beside `mix`, `load_slots`, `scratch`, `scratch_kb`; `HotClass { mb, k }`; names `hot32k4`, `hot64k2`; `Op::Hot`; the first `k` drawn load slots after the scratch slots are hot; no width roll for a hot class with version 2 widths (`takes_width_roll`); the program id carries `hot/S/k` |
|
|
| `igneum-pow/src/memhard.rs` | `HotTable` (key, S, words), `hot_key(seed_bytes)`, `hot_words(mb)`, `hot_segments(mb)`, `hot_index(src, words)`, the tagged segment fill shared with the cache, `HOT_TAG` |
|
|
| `igneum-pow/src/verify.rs` | `DatasetSource::hot: Option<HotTable>`; `Op::Hot` in `step`; `Epoch::new_class` and `from_seed_bytes_class` fill `H` from the program's seed bytes when the class has a hot table |
|
|
| `igneum-pow/src/accept.rs` | `Op::Hot` with the closed-form stand-in of 2.4 |
|
|
| `igneum-pow/src/emit.rs` | the hot load statement in the three dialects; `hot` as the buffer after the init words (Metal buffer 3, or 4 when bound; CUDA and OpenCL argument after `mask`, or after the init words when bound) and before the scratch triple; `HOT_WORDS` literal; `ht_cache_segment` and `igneum_hot_fill` kernels in memhard.h, kernel.cu, kernel.cl and memhard.metal; `program.h` `IGNEUM_HOT_MB`, `IGNEUM_HOT_WORDS`, `IGNEUM_HOT_SEGMENTS`, `IGNEUM_HOT_SLOTS`, `IGNEUM_HOT_KEY_INIT`; `vectors.h` and `vectors.json` the hot head, last line and FNV-1a 64 |
|
|
| `igneum-pow/src/main.rs` | `--class hot<S>k<k>` on every command (`bench` reports the fill time and the verifier ms per warp) |
|
|
| `igneum-pow/tests/packs.rs` | the five hot packs pinned (program, vectors, every emitted file, the load-form count: `16 - k` masked loads and `k` hot loads) |
|
|
| `proto-cuda/packs-ca2-hot/` | replaced: `hot32k4`, `hot64k4`, `hot96k4`, `hot64k2`, `hot64k8`; added: `hot32k4a`, `hot64k4a`, `hot96k4a`; all from seed `igneum-genesis`, day `2026-10-03` |
|
|
| `proto-cuda/nvrtc/packfile.h` | `hotMb`, `hotWords`, `hotSegments`, `hotSlots`; the hot self-test values; `pf_selftest` checks them when the pack carries them |
|
|
| `proto-opencl/host.c` | `--bench-pack` and the self-test allocate and fill `H` from the pack's `igneum_hot_fill`, check its head, last line and FNV-1a 64, pass it as the argument after the init words |
|
|
| `proto-metal/packbench.swift` | the same on Metal from `memhard.metal` |
|
|
| `proto-cuda/nvrtc/worker.cpp` | `--check` and `--bench` (new: the base kernel timed over `--batches` dispatches, one RESULT line) allocate, fill and self-test `H`; the launch passes it |
|
|
|
|
## 6. Measurements
|
|
|
|
Every row names the machine, the harness and the command; the bench-log entry of 5 October 2026 ("the hot table on the M5 Max") carries the raw lines. The Mac rows were taken on 5 October 2026, 20:19 to 20:21 UTC, under the measure lock, with the Mac's load average at 14 to 27 from other agents' CPU work (the lock serialises builds and measurements, not every process), so they are ordered, repeatable to within a few percent against each other, and not the Mac's quiet numbers. The PC rows wait for the coordinator's go.
|
|
|
|
### 6.1 Random-read probe at the hot sizes
|
|
|
|
`igneum-bench-cl-igneum-genesis-mh --memprobe --probe-mib S` (`proto-opencl/host.c`; dependent random 4-byte loads, best of 3, 256 steps per lane, work-group 256; the ceiling is the 4,194,304-lane row; Apple OpenCL reports wall time).
|
|
|
|
| Card | 32 MiB ceiling | 64 MiB | 96 MiB | 1024 MiB | Ratio 32 / 1024 | 64 / 1024 | 96 / 1024 | ns per dependent load at 4,096 lanes (32, 64, 96, 1024 MiB) |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| M5 Max (Apple OpenCL) | 21.7 G loads/s | 12.8 | 12.3 | 3.50 | 6.2 | 3.7 | 3.5 | 1,168; 1,129; 1,242; 1,844 |
|
|
| RTX 5090 (CUDA worker `--memprobe`, PC 1, job run-ca2-hot-5090-20261005, card off in the app) | 112.6 | 112.6 | 112.6 | 17.6 | 6.4 | 6.4 | 6.4 | 320; 340; 336; 610 |
|
|
| RX 9070 XT (OpenCL worker `--memprobe --device 1`, PC 1, job run-ca2-hot-9070-20261005, card off in the app) | 9.88 (10.8 at 262,144 lanes) | 9.47 | 8.18 | 2.43 | 4.1 | 3.9 | 3.4 | 396; 457; 454; 1,579 |
|
|
|
|
Reading, RTX 5090: all three sizes sit inside the 96 MiB L2 at one ceiling (112.6 G loads/s, 6.4x the DRAM figure), so the probe alone promises a full hit rate for every S. RX 9070 XT: 32 and 64 MiB inside the Infinity Cache at 9.5 to 10.8 G loads/s (3.9 to 4.1x), 96 MiB at 8.2 (3.4x), as the 9070 XT entry's 64 MiB row said. Streams: 5090 1,563 GB/s at 1024 MiB, 9070 XT 633 GB/s (both at their rated figures, the cards were not parked).
|
|
|
|
Reading, M5 Max: the step from 32 to 64 MiB halves the ceiling (21.7 to 12.8 G loads/s) and 96 MiB sits with 64, so 32 MiB is inside a cache level that 64 MiB is not (the M5 Max's system level cache size is not published; approximate reading: the 32 MiB table fits, the two larger ones mostly do not and run at a 3.5 to 3.7x advantage over DRAM from whatever hits they get). The coalesced stream grows with the buffer (139, 199, 243, 522 GB/s) because the small buffers are read once from cold.
|
|
|
|
Predicted `g` from the probe alone, with a hot load costing `1 / ceiling_S` and a cold load `1 / ceiling_1024`: `g = 1 / ((16 - k) / 16 + (k / 16) x ceiling_1024 / ceiling_S)`.
|
|
|
|
### 6.2 Bit-exactness and hash rate per hot pack
|
|
|
|
Metal: `proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24` (GPU time). Apple OpenCL: `igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 5 --batch-log2 24` (wall time). Both harnesses fill the hot table on the device from the pack's `igneum_hot_fill` and check its head, last line and FNV-1a 64 against vectors.json or vectors.h; vectors are the 96 lanes of the Rust reference; the fingerprint is FNV-1a 64 over the 2^24 outputs at base nonce 0.
|
|
|
|
| Pack | Vectors (Metal, OpenCL) | Hot table FNV (Metal, OpenCL) | Fingerprint 2^24 (both harnesses equal) | Metal Mhash/s | Apple OpenCL Mhash/s | g against v2 (Metal) | Predicted g from the probe | Ideal g |
|
|
|---|---|---|---|---|---|---|---|---|
|
|
| v2 (igneum-genesis-mh) | 3/3, 96/96 | none | 25f96e7dce90bd4e | 27.68 | 27.61 | 1 | 1 | 1 |
|
|
| hot32k4 | 3/3, 96/96 | PASS, PASS | d2e6cf3b61d0b9fe | 33.90 | 33.92 | 1.22 | 1.27 | 1.33 |
|
|
| hot64k4 | 3/3, 96/96 | PASS, PASS | e4c5263ac650cc0d | 30.93 | 30.89 | 1.12 | 1.22 | 1.33 |
|
|
| hot96k4 | 3/3, 96/96 | PASS, PASS | 5d63439b6e394521 | 29.06 | 28.97 | 1.05 | 1.22 | 1.33 |
|
|
| hot64k2 | 3/3, 96/96 | PASS, PASS | 352633bdbbb0d2b6 | 27.67 | 27.27 | 1.00 | 1.10 | 1.14 |
|
|
| hot64k8 | 3/3, 96/96 | PASS, PASS | da54630d7dfaaf85 | 47.42 | 46.73 | 1.71 | 1.57 | 2.0 |
|
|
| hot32k4a (added) | 3/3, 96/96 | PASS, PASS | 8a3414735db4523c | 25.76 | 25.72 | 0.93 | 0.96 | 1 |
|
|
| hot64k4a (added) | 3/3, 96/96 | PASS, PASS | 45668f34105f6307 | 23.92 | 23.87 | 0.87 | 0.94 | 1 |
|
|
| hot96k4a (added) | 3/3, 96/96 | PASS, PASS | af763997dfee4c82 | 22.92 | 22.88 | 0.83 | 0.93 | 1 |
|
|
|
|
For the added form the ideal `g` is 1 (the 16 dataset loads stay) and the probe predicts `g = 1 / (1 + (k / 16) x ceiling_1024 / ceiling_S)`: 0.96 at 32 MiB, 0.94 at 64 and 96 MiB for `k = 4` on the M5 Max; what matters is how far below 1 the GPU lands (its cost of the layer) against the chip's `S` MiB of SRAM and `k` reads.
|
|
|
|
The PCs (5 October 2026, 21:29 to 21:35 UTC, PC 1 ae432dc7 on app 0.3.9 before and after, the card under test switched off in the app through `api/cards` and restored; CUDA worker `--bench --batches 5 --batch-log2 24 --block-warps 1` on the RTX 5090, wall time around the stream sync; OpenCL worker `--bench-pack --batches 5 --batch-log2 24 --device 1` on the RX 9070 XT, device event time, work-group 256; the v2 references are the readwidth entry's same-night, same-worker numbers: 136.1 and 18.15 MH/s). Every pack bit-exact with the Mac's fingerprint, hot table head, last line and FNV PASS, 96/96 lanes, on both cards.
|
|
|
|
| Pack | RTX 5090 MH/s | g (v2 136.1) | probe-predicted g (5090) | RX 9070 XT MH/s | g (v2 18.15) | probe-predicted g (9070 XT) | ideal g |
|
|
|---|---|---|---|---|---|---|---|
|
|
| hot32k4 | 146.6 | 1.08 | 1.27 | 19.79 | 1.09 | 1.23 | 1.33 |
|
|
| hot64k4 | 140.8 | 1.03 | 1.27 | 18.73 | 1.03 | 1.23 | 1.33 |
|
|
| hot96k4 | 138.5 | 1.02 | 1.27 | 18.33 | 1.01 | 1.21 | 1.33 |
|
|
| hot64k2 | 137.5 | 1.01 | 1.13 | 18.17 | 1.00 | 1.11 | 1.14 |
|
|
| hot64k8 | 163.6 | 1.20 | 1.73 | 22.32 | 1.23 | 1.59 | 2.0 |
|
|
| hot32k4a (added) | 118.7 | 0.87 | 0.96 | 15.27 | 0.84 | 0.94 | 1 |
|
|
| hot64k4a (added) | 115.4 | 0.85 | 0.96 | 14.62 | 0.81 | 0.94 | 1 |
|
|
| hot96k4a (added) | 114.4 | 0.84 | 0.96 | 14.56 | 0.80 | 0.93 | 1 |
|
|
|
|
The 5090 at `--block-warps 8` (6 blocks per SM, 8,160 resident warps) is within 0.7% of every row above; the 9070 XT at work-group 32 within 0.5%. Hot fill on the 5090: 1.2 to 2.5 ms (wall, driver API); on the 9070 XT 1.6 to 5.3 ms.
|
|
|
|
Reading, the PCs. The probe promised a full hit rate for every S on the 5090 (all three tables inside the 96 MiB L2 at one ceiling) and 3.4 to 4.1x on the 9070 XT, and the hash delivered a fraction of it: 1.02 to 1.08x at `k = 4` against the probe's 1.27x and the ideal 1.33x on the 5090, 1.01 to 1.09x on the 9070 XT, 1.20 to 1.23x at `k = 8` against 1.73 and 2.0x. The table that stands alone in the probe does not stand up with the 1 GiB dataset streaming through the same cache: the dataset's random lines evict it (the 5090's L2 and the 9070 XT's Infinity Cache are shared by every load; neither card partitions them). The added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% of its rate for four extra loads per iteration, three to five times the probe's 4 to 7%. The three cards agree on the shape; the 5090's larger cache buys it nothing over the Mac at 32 MiB (1.08 against 1.22).
|
|
|
|
Hot table fill on the M5 Max GPU (Metal, GPU time): 0.07 ms at 32 MiB, 0.15 ms at 64 MiB, 0.22 ms at 96 MiB.
|
|
|
|
Reading, M5 Max. Bit-exactness holds: Metal and Apple OpenCL give one fingerprint per pack and every vector lane against the Rust reference, hot table included. On the rate, the hot loads are worth 92% of the probe's prediction at 32 MiB (1.22 against 1.27), 50% at 64 MiB (1.12 against 1.22 is 0.12 of 0.22) and 23% at 96 MiB; at `k = 8` the measured 1.71 is above the prediction (1.57), which says the cold loads also go faster when half of them are gone (fewer dependent DRAM reads in the chain per hash). `hot64k2` gained nothing on this Mac at load (27.67 against 27.68). So on Apple silicon, with the 1 GiB dataset streaming through the same cache, the table that fits (32 MiB) delivers most of its ideal and the larger ones lose most of theirs. The 5090 and 9070 XT, whose caches are the plan's targets, are the PC job.
|
|
|
|
### 6.3 CPU verifier, one core, avg of 50 warps
|
|
|
|
`igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <class> --warps 50` (release build, one M5 Max core, the Mac at load as above; the v2 row from the same session, the 0.604 ms of the brief was an earlier quiet run).
|
|
|
|
| Class | Hot fill, one core | Items derived per warp | ms per warp | Against v2 |
|
|
|---|---|---|---|---|
|
|
| v2 | none | 4,096 | 0.626 | 1 |
|
|
| hot32k4 | 24.0 ms | 3,072 | 0.489 | 0.78 |
|
|
| hot64k4 | 46.4 ms | 3,072 | 0.504 | 0.81 |
|
|
| hot96k4 | 73.0 ms | 3,072 | 0.488 | 0.78 |
|
|
| hot64k2 | 45.5 ms | 3,584 | 0.560 | 0.89 |
|
|
| hot64k8 | 47.7 ms | 2,048 | 0.344 | 0.55 |
|
|
| hot32k4a (added) | 21.7 ms | 4,096 | 0.631 | 1.05 (v2 in the same session 0.602) |
|
|
| hot64k4a (added) | 43.3 ms | 4,096 | 0.609 | 1.01 |
|
|
| hot96k4a (added) | 64.9 ms | 4,096 | 0.614 | 1.02 |
|
|
|
|
Reading: the verifier gets cheaper with `k`, because a hot load is one table read where a dataset load is an item derivation (9 mixers and 8 cache reads); the per-epoch cost is the fill, 24 to 73 ms on one core, against a 3,600 s epoch and the 20-minute seed lead. The 10 ms gate (spec 1.16) is unaffected.
|
|
|
|
### 6.4 The chip model with the measured g (M5 Max; the PC rows will replace it)
|
|
|
|
Chip A (no SRAM for `H`): gain after = `G0 / g`. Chip B (`H` in SRAM): `G0 x A_die / (A_die + A_S)` at 0.39 mm^2 per MiB, 100 mm^2 die.
|
|
|
|
| Class | g (M5 Max) | Chip A, gain after as a share of G0 | Chip B, share of G0 at 100 mm^2 | Chip B at 750 mm^2 |
|
|
|---|---|---|---|---|
|
|
| hot32k4 | 1.22 | 0.82 | 0.89 | 0.98 |
|
|
| hot64k4 | 1.12 | 0.89 | 0.80 | 0.97 |
|
|
| hot96k4 | 1.05 | 0.95 | 0.73 | 0.95 |
|
|
| hot64k8 | 1.71 | 0.58 | 0.80 | 0.97 |
|
|
|
|
The added form (second Mac session, 21:03 to 21:19 UTC, load average 7 to 14; v2 in that session 27.63 Metal, 27.59 OpenCL): the GPU keeps 0.93 / 0.87 / 0.83 of its rate at 32 / 64 / 96 MiB (`k = 4`), against the probe's 0.96 / 0.94 / 0.93, so the hot hits cost this card more than the probe says as the table grows, and 96 MiB is not resident. The named chip's position under the added form, same assumptions:
|
|
|
|
| Chip | Hot loads served from | Rate against its own version 2 rate | Gain after, as a share of G0 (S = 32 / 64 / 96) | Arithmetic |
|
|
|---|---|---|---|---|
|
|
| A, no SRAM for H | DRAM | at most 16 / 20 = 0.80 (20 dependent DRAM loads per iteration in place of 16) | 0.86 / 0.92 / 0.96 | `0.80 / g` |
|
|
| B, S MiB of SRAM for H, 100 mm^2 die, 0.39 mm^2/MiB | SRAM (the k reads near free) | 1.00 | 0.96 / 0.92 / 0.88 | `(1 / g) x A_die / (A_die + A_S)` |
|
|
| B on a 750 mm^2 die | SRAM | 1.00 | 1.06 / 1.12 / 1.15 | the SRAM is 1.6 to 4.8% of the die, the GPU's loss is larger |
|
|
|
|
With the PCs' `g` (added form, `k = 4`): RTX 5090 0.87 / 0.85 / 0.84, RX 9070 XT 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB.
|
|
|
|
| Chip, added form, S = 32 / 64 / 96 MiB | Gain after as a share of G0, 5090 g | 9070 XT g |
|
|
|---|---|---|
|
|
| A, H from DRAM (rate 0.80 of its own) | 0.92 / 0.94 / 0.95 | 0.95 / 0.99 / 1.00 |
|
|
| B, S MiB of SRAM, 100 mm^2 die, 0.39 mm^2/MiB | 1.03 / 0.94 / 0.87 | 1.06 / 0.99 / 0.91 |
|
|
| B, 750 mm^2 die | 1.13 / 1.14 / 1.17 | 1.17 / 1.20 / 1.22 |
|
|
|
|
Reading: on the cards the plan named, the added form taxes no chip. A chip that serves H from its DRAM loses at most 8% of its gain on the 5090's figures and nothing on the 9070 XT's; a chip with the SRAM comes out ahead on every row except the smallest die at 64 and 96 MiB. The replaced form helps the recompute chip outright (section 2.3). The hot hits are not free on a GPU while the 1 GiB dataset streams through the same cache, and the layer's premise (section 1) needed them to be.
|
|
|
|
**Recommendation (owner: the project lead, gate 1):** do not adopt layer 5 in either form on these measurements. What could change it: a GPU-side way to keep H resident (cache partitioning or persisting-access controls exist on NVIDIA, approximate, from memory, and are a driver setting, not a consensus rule), or a dataset access pattern that bypasses the cache; both are outside the hash and were not measured. The experiment stays behind the flag with its packs and vectors as the record.
|
|
|
|
Reading, replaced form: on the M5 Max figures the chip's cheaper way out of `hot64k8` is the SRAM (0.80 of `G0` at a 100 mm^2 die) rather than serving `H` from DRAM (0.58). The layer's value is then the SRAM area, which is small on a large die (0.97 at 750 mm^2). The strongest configuration on this card is the one whose table fits the cache and whose `k` is large; whether 64 MiB fits the 5090's L2 and the 9070 XT's Infinity Cache with the dataset streaming beside it is the PC measurement.
|
|
|
|
## 7. Decision and what would make it pay (the Counter ASIC 3.0 note)
|
|
|
|
Decided by the coordinator on 5 October 2026 after the PC rows: layer 5 is out of v3. The rule was a GPU cost under 3% (`g` above 0.97) for a chip cost worth having; measured `g` 0.84 to 0.87 on the RTX 5090 and 0.80 to 0.84 on the RX 9070 XT for the added form, and the replaced form helps the recompute chip. No re-run.
|
|
|
|
What would make a hot table pay, for a later round:
|
|
|
|
| Condition | What the measurements say | What would have to be shown |
|
|
|---|---|---|
|
|
| A table that stays resident beside a streaming 1 GiB | The probe's ceiling at every S inside the cache (5090: 112.6 G loads/s at 32, 64 and 96 MiB) and the hash's 1.02 to 1.08x say the dataset's random lines evict it; 32 MiB on the M5 Max kept 92% of its probe gain, the 5090 kept 29% | A size small enough to survive: the probe cannot say which (it has no competing stream); a sweep of `S` down from 32 MiB (16, 8, 4, 2) with the hash itself, on the 5090 first, would. A 2 to 4 MiB table costs a chip under 2 mm^2 of SRAM, so the layer would then tax nothing; the point of the layer was a table a chip cannot afford, and a table a GPU keeps is one a chip affords |
|
|
| The `k` and `S` the probe says | `g` moved with `k` (1.20x at `k = 8`, 1.02 to 1.08 at `k = 4`, 1.00 at `k = 2`) and barely with `S`; the probe predicted 1.27 to 1.73 | Only a resident table makes `k` worth raising; at `k = 8` with a resident table the replaced form reaches 2x (ideal) and the added form costs 0 by the probe. Without residency, `k` buys rate on the replaced form (which helps the chip) and costs rate on the added form |
|
|
| A different access shape | Both forms read H at one random word per hot load, the dataset's own pattern, and share the cache with 128 dataset reads per hash | A shape that touches the table in cache-line units and few lines per hash (one 64-byte line per iteration, say) would need a smaller resident set per hash and could be measured with the read-width emitters (`wide_load_stmt`) pointed at H. Or a hot table whose lines are read in a fixed order per epoch (a stream, not a random read), which a GPU prefetches and a chip must still hold or fetch; not designed here |
|
|
| Cache partitioning on the GPU | NVIDIA exposes persisting L2 access controls to CUDA programs (approximate, from memory); AMD and Apple do not expose an equivalent to OpenCL or Metal | A consensus rule cannot depend on a driver feature of one vendor; it could be a miner-side optimisation if the layer were in, which it is not |
|
|
|
|
What the round leaves in place: the code behind the flag (`LoadClass::hot`, both forms, the hosts filling H on the device from the epoch seed, the eight packs and their vectors), the probe at the hot sizes on three cards, and the measured rule that a read-only table a GPU cache could hold is not one it does hold while the dataset streams.
|
|
|
|
## 8. What is unverified
|
|
|
|
Listed here until measured, and carried into the bench-log entry.
|
|
|
|
1. Measured on all three cards (6.2): no card keeps `H` resident enough to deliver the probe's hit rate while the dataset streams. Not measured: whether a driver-side cache partition (NVIDIA's persisting L2 access controls, approximate, from memory) would; it is not something a consensus rule can rely on.
|
|
2. The PC rows are one run each (jobs run-ca2-hot-5090-20261005 and run-ca2-hot-9070-20261005, 5 October 2026); the two launch shapes per card agree within 1%, and no re-run was taken (PC 1 was handed on).
|
|
3. The resident warp counts of section 3 are approximate; the occupancy the hosts reach is printed by each harness (`kernel:` lines) and should replace them.
|
|
4. The SRAM area figures are approximate (section 4 cites their sources); no chip was priced.
|
|
5. `S = 96` uses the multiply-shift mapping; the spec's dataset still uses `AND MASK`. If the layer goes in, gate 1 decides whether the dataset mapping follows (section 1.13.3) or `S` stays a power of two.
|
|
6. The hot table on the chain needs the epoch seed about 1 ms (GPU) or up to 71 ms (CPU) before the epoch starts; the 20-minute lead of section 4.3 covers it. Not exercised on a node.
|
|
7. The Mac numbers were taken at load average 14 to 27 (other agents' CPU work); the ratios between packs of one session are the result, the absolute Mhash/s are not quiet numbers.
|