The Swift DatasetContext is the version 2 item construction; a v3 program over it would hash another dataset than
the node's (every found refused by the CPU re-check). servePackDataset compiles the pack's memhard.metal and runs
igneum_cache_fill and igneum_build as packbench does, releases the cache, and the v3 job path takes the pack
program and the pack day of its class and era only (need + error otherwise). The gate script's --metal mode drives
igneum-miner --worker <igneum-bench> --prepare-packs <dir> --exit-on-seed-change (the app's shape) on node 0 and
reports the PREPARE and prepared lines, need / mismatch / refusal lines, accepted blocks after the switch, the
CPU re-check counter and the swap line.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.
Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
main.swift: servePackProgram reads program.h with the packfile.h checks (generator 2 or 3, the class line against
the generator, the seed bytes, IGNEUM_SEEDW_INIT against attempt_words, class and era against the line) and
compiles program_bound.metal; the program store keys on (seed, class, era); a v3 job with no resident v3 pack
program answers need + error; v2 lines unchanged (Swift generation, the variant race); a pack program never races.
verify.rs: Epoch::chain_dataset_day(day, class, days_since_genesis, genesis_dataset_log2) and days_since_genesis,
the entry the node builds every day cache through (the ca2-mixer growth rule fills the body).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
packfile.h: pf_load refuses a generator other than 2 or 3 (spec 01 section 1.4.5), reads IGNEUM_PROGRAM_CLASS
(must match the generator) and IGNEUM_ERA_SEED_HEX; pf_pack_class_ok and pf_class_token are the one rule for
the `class=<v2|v3> era=<hex>` tokens a job or prepare line of a class v3 epoch ends with (a v2 line is the
line of before, byte for byte). CUDA worker.cpp and OpenCL host.c: a pair's identity includes its class and era
when the line names them, so a prepared pack of the right class wins over the resident pack of the same seeds;
right seeds with the wrong class answer `need` plus `error <id> pack <dir>: program class mismatch ...`, and a
prepare on a pack of the wrong class fails in plain words. Metal main.swift: refuses every class v3 line until
the Swift generator carries version 3 (the integration branch). emu/packfile-test.c: the generator rule, a v3
pack with its era, the token matcher, the contradicting class line (13 checks pass).
igneum-pow verify.rs: Epoch::chain_dataset(day, class), the one entry the node's day cache goes through, so the
ca2-mixer class-specific item construction and cache schedule have a seam.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline
built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs
on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple;
2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's
vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two
seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped,
always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so
far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time.
A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model.
NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the
pack format, the miner and igneum-pow are untouched and old packs race. --race --pack <dir> runs the race alone.
Under IGNEUM_EMU the race is off (the stand-in checks the handed-over text is the pack's). Metal: the MSL hooks in
generateMSL, a lock-protected kernel slot per program, a pair compiled inline races after its first job,
--race-test --seed --day runs the race alone with a table. OpenCL is not raced yet. docs/design/miner-tuning.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a
load's source from the registers written earlier and not read by a load since, the other
48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no
stale load source, every register injected; dynamic: 64 units on the seed-keyed
closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias
within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced
by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the
generator version, attempt and program id. Version 1 kept as generate_v1 for the census.
Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export;
new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of
39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical
programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000);
CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs
at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887.
Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA,
OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
proto-metal --serve compiles igneum_hash_bound at runtime and mines jobs from stdin (init words in buffer 3, program
and dataset cached per seed); proto-cuda and proto-opencl host serve modes from the pack's kernel_bound.cu / .cl with a
seed guard; --vendor device filter for OpenCL. windows-miner/: START-MINING.bat + start-mining.ps1 (GPU and tool
detection, pack export, cached builds, MINERS identities per vendor, status every 30 s, uploads every 60 s, rebuild on
seed change, Ctrl+C summary), README.txt, make-package.sh (igneum-mine-test.zip with a cross-compiled miner).
docs: fork-divergence devnet v1, bench-log entry with the CPU, Metal and overnight numbers.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.
proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.
Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
proto-metal: default dataset is now the memory-hard construction (MEMHARD.md), --closed-form keeps the original.
Cache fill 2 ms GPU / 185 ms one CPU core; dataset build 20.6 ms; GPU cache == CPU cache on all 2^26 words.
Shortcut ratio: inline kernel 111x faster than honest (closed form) to 4.8x slower (memory-hard), 1 GiB.
CPU verify 0.63 to 0.80 ms per warp at 104 loads, 1.21 ms at 144 loads (4,608 items): 10 ms gate met.
Levers --load-weight and --wide-frac implemented and measured, both off; default generator unchanged.
Fuzz 200/200, edge, determinism, memcheck, stats re-run on the new dataset, all PASS.
proto-cuda: host.cu handles both dataset modes; new pack igneum-genesis-mh with memhard.h; clang emulation PASS
including the three-way cache check. Old packs unchanged; closed-form export is byte-identical to them.
docs/bench-log.md: dated summary.
Note: a concurrent session running git commit -a swept earlier states of these files into its site commits
(7b28d5e through d6539fa); this commit carries the remainder.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>