LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.
Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
packfile.h: pf_load refuses a generator other than 2 or 3 (spec 01 section 1.4.5), reads IGNEUM_PROGRAM_CLASS
(must match the generator) and IGNEUM_ERA_SEED_HEX; pf_pack_class_ok and pf_class_token are the one rule for
the `class=<v2|v3> era=<hex>` tokens a job or prepare line of a class v3 epoch ends with (a v2 line is the
line of before, byte for byte). CUDA worker.cpp and OpenCL host.c: a pair's identity includes its class and era
when the line names them, so a prepared pack of the right class wins over the resident pack of the same seeds;
right seeds with the wrong class answer `need` plus `error <id> pack <dir>: program class mismatch ...`, and a
prepare on a pack of the wrong class fails in plain words. Metal main.swift: refuses every class v3 line until
the Swift generator carries version 3 (the integration branch). emu/packfile-test.c: the generator rule, a v3
pack with its era, the token matcher, the contradicting class line (13 checks pass).
igneum-pow verify.rs: Epoch::chain_dataset(day, class), the one entry the node's day cache goes through, so the
ca2-mixer class-specific item construction and cache schedule have a seam.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Nothing changes for the default class: the pinned packs are byte-identical (tests/packs.rs), the v2 draw stream is untouched.
LoadClass {mix, load_slots, scratch}: fixed widths w16, w64, w64x4 (4 loads of 64 B), era mixes 50/35/15 and 25/50/25 drawn per load with one extra below(100) roll, and the scratch variant scr0/2/4/8 (persistent warps, 1 MiB per warp, tagged lazy fill, measurement only). A wide load reads the W-aligned address and folds every word: x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]. Program ids carry the class. proto-opencl/host.c taken from opencl-rdna4 23810df (--memprobe, select read-back).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
From 18:23Z both Windows workers (CUDA on PC 1 and PC 2, OpenCL on PC 1 after the 18:34Z node restart)
refused every pack for epoch 34 with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", and
the miner and the app restarted them every 5 to 60 s until 18:44Z and beyond. The packs were correct.
The generator retries a rejected candidate with seed || k_le32 (attempt_words); epoch 34's attempt 0 was
rejected (246 of 16384 final register values saturated, limit 163) and attempt 1 accepted, so the pack
carried attempt 1's words while pf_load (proto-cuda/nvrtc/packfile.h) derived the expected words from
the bare seed. Both workers also matched jobs to pairs by those bare-seed words, so even a loaded pack
of a retried program would have answered "epoch seed mismatch" on every job.
The rule, in one place per language:
- packfile.h: pf_program_words(bytes, attempt); pf_load reads IGNEUM_PROGRAM_ATTEMPT and checks the
attempt's words; the refusal says "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not
attempt N of the epoch seed ..." in plain words.
- worker.cpp and proto-opencl/host.c: a job belongs to a pair when the seed hex the node sent is the
pair's (pairIs); the compiled-in placeholder pack keeps the word comparison.
- igneum-pow/src/packcheck.rs: verify_pack_texts / verify_pack_dir, the same rule in Rust; the miner
checks every pack it writes with it before a worker sees it (vendor/igneum-node pack-loop branch).
Tests pin the attempt vectors of epoch 34 on both sides (one vector, two implementations), that
epoch 34 is attempt 1 and epoch 33 attempt 0, a known-good pack of a later attempt, a known-mismatched
(out of date) pack, and self-contradicting packs.
- proto-cuda/nvrtc/emu/packfile-test.c (+ .sh, in CI): pf_load on a known-good attempt-1 pack, the
checked-in attempt-0 pack, and the known-mismatched bare-words pack.
The app (app/igneum-app):
- watchdog.rs: PACK_OUT_OF_DATE_CODE 44, PackRebuilds (at most 3 pack exports per epoch, then the card
shows the reason), pack_refusal (the worker's "error 0 pack" line and the miner's "PACK OUT OF DATE"
line), pack_epoch_of; tests on the incident lines, known-good and known-mismatched.
- engine.rs: exit 44 exports the pack again before the restart instead of a blind restart, the strip
says "program pack out of date, rebuilding", the card and the log name the condition; at the cap the
card is marked failed with the reason and tries again in 10 minutes.
The relay (relay/lib/parse.mjs): PACK_MISMATCH; the card reads "pack mismatch, rebuilding (N refusals
in the tail, M restarts)" in `node tools/console.mjs machines` instead of a bare restart count; tests
on the PC 2 tail of 18:27Z and a healthy tail.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PC 1 after Adrenalin 26.9.2 (5 October 2026): two AMD ICDs each listed the iGPU and the RX 9070 XT, so the 0.3.9
app showed five rows for three GPUs (gfx1036, gfx1201, gfx1036, gfx1201 plus the 5090), ran two workers on the one
9070 XT at 9 MH/s each, called the iGPU discrete, and re-enabled it when its key moved from amd:1:gfx1036 to
amd:0:gfx1036.
detect.rs: the Windows path is a pure assemble(Inputs) over nvidia-smi, the worker's --list and the adapter list.
dedupe_platforms keeps one entry per card per vendor: the fuller platform wins (most GPUs, then the newest driver,
then the first listed) and another platform's device survives only at a PCI address the winner lacks (without
addresses: the same code and ordinal); the dropped entries go to the log, never to a worker. resolve_name puts the
Windows adapter name on the row: by PCI bus (the adapter list now carries DEVPKEY_Device_BusNumber and Address),
else by the gfx code's PCI device ids, else by the card names in a gfx table (gfx1201, 1200, 1100, 1101, 1102,
1030, 1031, 1032, 1036, 1035, 1103, 1150, 90c, approximate), else the table's words, else the code. The code stays
in `code`, the key and the row's tooltip; the key is vendor:code with "#2" for a twin, no index. hotplug::pref_for
reads a 0.3.9 key (vendor:index:code) when exactly one matches; the diff matches by key and PCI address, then by
vendor and name or code. The iGPU is integrated and off by default; the choice survives a moved index.
proto-opencl/host.c: --list prints ", pci bb:dd.f" on the info line from CL_DEVICE_TOPOLOGY_AMD or the NVIDIA bus
and slot ids (needs a worker rebuild through proto-cuda/nvrtc/build-windows.sh and push-inputs; an old worker
still works, by code and ordinal).
Tests: detect::tests has the five-row PC 1 list as a fixture (two platforms, the three adapters with DEV_13C0,
DEV_2B85 and DEV_7550) and asserts three rows named NVIDIA GeForce RTX 5090, AMD Radeon(TM) Graphics (integrated,
off) and AMD Radeon RX 9070 XT (discrete, on, 8 identities), the keys, the console line, and the diff against the
one-platform list before the eGPU (iGPU moved, 9070 XT added, nothing removed); plus the list parser with and
without pci, dedupe with and without addresses, name resolution, twins' keys. cargo test -p igneum-app: 96 passed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
After PC 2's gfx1036 (prebuilt-generic path) completed 2,000 jobs a second with no hash from 600 s on: every OpenCL
call in host.c's job path is now fatal on error (exit 3, the miner restarts the worker), the dispatch event must read
CL_COMPLETE, a chunk 20x faster per nonce than the running mean or an output buffer unchanged since the previous
dispatch is a fault, and a stats line every 200 jobs carries the live event and buffer counts (a leak over 16 events
or 12 buffers is fatal too). IGNEUM_FAULT_TEST=N exercises the detectors on a healthy device (verified on Apple
OpenCL: the stale-output guard fires on the chunk after the injected fault). The launcher shows 'worker fault' and
'restarting' for the card on the miner's WORKER FAULT line and drops the last rate. The emulation's NVRTC stand-in
now applies NVRTC's execution-space rule (program.h(46) igneum_launch_* declarations are host code unless
-default-device or -DIGNEUM_NO_CUDA), which is what the RTX 5090 reported; test.sh checks the rejection.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
host.c --serve --pack <dir> serves a pack read at run time (packfile.h), whatever pack the exe was built against: the
first pair is built and self-tested like a prepared one (pairBuffers, pairSelfTest), prepared pairs are self-tested
too, the ready line says path prebuilt-generic. cl_dynamic.h (IGNEUM_CL_DYNAMIC) opens OpenCL.dll with LoadLibrary and
fills a function pointer per entry point, so the mingw build links no import library. test-generic.sh checks the mode
through Apple OpenCL with the packs of the CUDA emulation test: PASS (192 found, prepare and swap, 15 sampled hashes
equal igneum-pow hash-bound), built against a different placeholder pack than it served.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
proto-metal --serve compiles igneum_hash_bound at runtime and mines jobs from stdin (init words in buffer 3, program
and dataset cached per seed); proto-cuda and proto-opencl host serve modes from the pack's kernel_bound.cu / .cl with a
seed guard; --vendor device filter for OpenCL. windows-miner/: START-MINING.bat + start-mining.ps1 (GPU and tool
detection, pack export, cached builds, MINERS identities per vendor, status every 30 s, uploads every 60 s, rebuild on
seed change, Ctrl+C summary), README.txt, make-package.sh (igneum-mine-test.zip with a cross-compiled miner).
docs: fork-divergence devnet v1, bench-log entry with the CPU, Metal and overnight numbers.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.
proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.
Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>