Intel's compiler turns rotate(x, (0u - n) & 31u) into a rotate LEFT by n: lane 0's register trace on the B580 diverged
at instruction 6 of iteration 0 (rotr) and nowhere before, in both exchange modes, with every other family and the
dataset kernels bit-exact. proto-opencl/intel_rotr.h rewrites the one helper line when the device's vendor or
platform string holds Intel (host.c's buildProgram and the prepare path), no other vendor sees a change, no pack or
consensus text moves. proto-opencl/test_intel_rotr.c (the pre-push gate runs it) feeds the line through the rewrite
under Intel, AMD and NVIDIA strings and asserts the outputs.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 26e135a362718a68a842da59080b43e93afd9dc2)
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.
Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
the project lead watched the 9070 XT at 90% usage with its fans barely turning and the app could not say what it drew: the
draw, temperature and MH per watt line came from nvidia-smi only, and the earlier per-watt figure used the board
rating. proto-opencl/gpu-telemetry.c prints one line per AMD card per sample (bus from SetupAPI by the display
device's name, kind, name, watts, temp_c, fan_rpm, fan_pct, mclk_mhz, gclk_mhz, util_pct, source), built by
build-windows.sh against vendor/adlx (the SDK clone), shipped by make-payload.sh and push-inputs.sh. The engine
runs it with -l 5 beside nvidia-smi (Source::AmdTelemetry, tick_amd_telemetry), parse_amd_telemetry fills
power_w, temp_gpu, fan_pct, fan_rpm, mclk_mhz, util_pct and telemetry_at on the AMD card matched by kind and
ordinal, so eff_mhw and the dashboard's existing line show it; app.js shows fan and memory clock when present.
Tests: three on the parser with lines captured on PC 1 and the Mac fixture; the sysfs path ran on a fixture tree.
Measured over 20:27:45 to 20:29:41 UTC with both cards mining (docs/bench-log.md, under the 9070 XT ceiling table):
9070 XT 198.9 W (193 to 212), 64 C, 657 rpm, 2,505 MHz memory, 3,290 MHz shader, 100% busy, 17.73 MH/s =
0.089 MH/W; RTX 5090 307.6 W, 69 C, 44% fan, 122.30 MH/s = 0.398 MH/W.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
packfile.h: pf_load refuses a generator other than 2 or 3 (spec 01 section 1.4.5), reads IGNEUM_PROGRAM_CLASS
(must match the generator) and IGNEUM_ERA_SEED_HEX; pf_pack_class_ok and pf_class_token are the one rule for
the `class=<v2|v3> era=<hex>` tokens a job or prepare line of a class v3 epoch ends with (a v2 line is the
line of before, byte for byte). CUDA worker.cpp and OpenCL host.c: a pair's identity includes its class and era
when the line names them, so a prepared pack of the right class wins over the resident pack of the same seeds;
right seeds with the wrong class answer `need` plus `error <id> pack <dir>: program class mismatch ...`, and a
prepare on a pack of the wrong class fails in plain words. Metal main.swift: refuses every class v3 line until
the Swift generator carries version 3 (the integration branch). emu/packfile-test.c: the generator rule, a v3
pack with its era, the token matcher, the contradicting class line (13 checks pass).
igneum-pow verify.rs: Epoch::chain_dataset(day, class), the one entry the node's day cache goes through, so the
ca2-mixer class-specific item construction and cache schedule have a seam.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Nothing changes for the default class: the pinned packs are byte-identical (tests/packs.rs), the v2 draw stream is untouched.
LoadClass {mix, load_slots, scratch}: fixed widths w16, w64, w64x4 (4 loads of 64 B), era mixes 50/35/15 and 25/50/25 drawn per load with one extra below(100) roll, and the scratch variant scr0/2/4/8 (persistent warps, 1 MiB per warp, tagged lazy fill, measurement only). A wide load reads the W-aligned address and folds every word: x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]. Program ids carry the class. proto-opencl/host.c taken from opencl-rdna4 23810df (--memprobe, select read-back).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
From 18:23Z both Windows workers (CUDA on PC 1 and PC 2, OpenCL on PC 1 after the 18:34Z node restart)
refused every pack for epoch 34 with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", and
the miner and the app restarted them every 5 to 60 s until 18:44Z and beyond. The packs were correct.
The generator retries a rejected candidate with seed || k_le32 (attempt_words); epoch 34's attempt 0 was
rejected (246 of 16384 final register values saturated, limit 163) and attempt 1 accepted, so the pack
carried attempt 1's words while pf_load (proto-cuda/nvrtc/packfile.h) derived the expected words from
the bare seed. Both workers also matched jobs to pairs by those bare-seed words, so even a loaded pack
of a retried program would have answered "epoch seed mismatch" on every job.
The rule, in one place per language:
- packfile.h: pf_program_words(bytes, attempt); pf_load reads IGNEUM_PROGRAM_ATTEMPT and checks the
attempt's words; the refusal says "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not
attempt N of the epoch seed ..." in plain words.
- worker.cpp and proto-opencl/host.c: a job belongs to a pair when the seed hex the node sent is the
pair's (pairIs); the compiled-in placeholder pack keeps the word comparison.
- igneum-pow/src/packcheck.rs: verify_pack_texts / verify_pack_dir, the same rule in Rust; the miner
checks every pack it writes with it before a worker sees it (vendor/igneum-node pack-loop branch).
Tests pin the attempt vectors of epoch 34 on both sides (one vector, two implementations), that
epoch 34 is attempt 1 and epoch 33 attempt 0, a known-good pack of a later attempt, a known-mismatched
(out of date) pack, and self-contradicting packs.
- proto-cuda/nvrtc/emu/packfile-test.c (+ .sh, in CI): pf_load on a known-good attempt-1 pack, the
checked-in attempt-0 pack, and the known-mismatched bare-words pack.
The app (app/igneum-app):
- watchdog.rs: PACK_OUT_OF_DATE_CODE 44, PackRebuilds (at most 3 pack exports per epoch, then the card
shows the reason), pack_refusal (the worker's "error 0 pack" line and the miner's "PACK OUT OF DATE"
line), pack_epoch_of; tests on the incident lines, known-good and known-mismatched.
- engine.rs: exit 44 exports the pack again before the restart instead of a blind restart, the strip
says "program pack out of date, rebuilding", the card and the log name the condition; at the cap the
card is marked failed with the reason and tries again in 10 minutes.
The relay (relay/lib/parse.mjs): PACK_MISMATCH; the card reads "pack mismatch, rebuilding (N refusals
in the tail, M restarts)" in `node tools/console.mjs machines` instead of a bare restart count; tests
on the PC 2 tail of 18:27Z and a healthy tail.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PC 1 after Adrenalin 26.9.2 (5 October 2026): two AMD ICDs each listed the iGPU and the RX 9070 XT, so the 0.3.9
app showed five rows for three GPUs (gfx1036, gfx1201, gfx1036, gfx1201 plus the 5090), ran two workers on the one
9070 XT at 9 MH/s each, called the iGPU discrete, and re-enabled it when its key moved from amd:1:gfx1036 to
amd:0:gfx1036.
detect.rs: the Windows path is a pure assemble(Inputs) over nvidia-smi, the worker's --list and the adapter list.
dedupe_platforms keeps one entry per card per vendor: the fuller platform wins (most GPUs, then the newest driver,
then the first listed) and another platform's device survives only at a PCI address the winner lacks (without
addresses: the same code and ordinal); the dropped entries go to the log, never to a worker. resolve_name puts the
Windows adapter name on the row: by PCI bus (the adapter list now carries DEVPKEY_Device_BusNumber and Address),
else by the gfx code's PCI device ids, else by the card names in a gfx table (gfx1201, 1200, 1100, 1101, 1102,
1030, 1031, 1032, 1036, 1035, 1103, 1150, 90c, approximate), else the table's words, else the code. The code stays
in `code`, the key and the row's tooltip; the key is vendor:code with "#2" for a twin, no index. hotplug::pref_for
reads a 0.3.9 key (vendor:index:code) when exactly one matches; the diff matches by key and PCI address, then by
vendor and name or code. The iGPU is integrated and off by default; the choice survives a moved index.
proto-opencl/host.c: --list prints ", pci bb:dd.f" on the info line from CL_DEVICE_TOPOLOGY_AMD or the NVIDIA bus
and slot ids (needs a worker rebuild through proto-cuda/nvrtc/build-windows.sh and push-inputs; an old worker
still works, by code and ordinal).
Tests: detect::tests has the five-row PC 1 list as a fixture (two platforms, the three adapters with DEV_13C0,
DEV_2B85 and DEV_7550) and asserts three rows named NVIDIA GeForce RTX 5090, AMD Radeon(TM) Graphics (integrated,
off) and AMD Radeon RX 9070 XT (discrete, on, 8 identities), the keys, the console line, and the diff against the
one-platform list before the eGPU (iGPU moved, 9070 XT added, nothing removed); plus the list parser with and
without pci, dedupe with and without addresses, name resolution, twins' keys. cargo test -p igneum-app: 96 passed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Decision of 4 October 2026 (evening): the Igneum Miner software takes a visible, switchable 1% dev fee, the norm
for GPU miners; the protocol stays fee-free. The miner side (--dev-fee, the 1-in-100 template counter, the audit
command) is on branch dev-fee of the node fork.
- app: settings.dev_fee (default on) passes --dev-fee 0 to igneum-miner when off; Settings shows the miner's own
"dev fee 1% (1 block in 100) to 0x..." line next to the rewards address with a switch; the engine parses the
miner's start line and its dev-fee block lines (fee_session, lifetime fee_total, an event per fee block)
- packaging/hive: h-manifest.conf, h-config.sh, h-run.sh, h-stats.sh, make-hive-package.sh (igneum-hive-<v>.tar.gz
with the Linux igneumd, igneum-miner and both GPU workers), README with the Flight Sheet, selftest.sh (bash -n,
stub binaries, the three hooks the way Hive runs them, the stats JSON parsed). Hive itself is untested
- infra/cross/build-workers-linux.sh: the NVRTC and OpenCL workers cross-compiled for Linux with zig;
proto-opencl/cl_dynamic.h gains the Linux dlopen branch (libOpenCL.so.1)
- tools/dev-fee/run.mjs: the fee-block test network (two nodes on 29900+, three CPU miners, payouts audit)
- docs/design/miner-dev-fee.md (mechanism, flag, lines, the DEV_FEE_ADDRESS placeholder and the devnet address),
docs/fud-ledger.md E18 and the E5/L9 status line, litepaper "What a miner's hour looks like" paragraph, homepage
miner section note
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every Windows executable now ships with the coin icon Explorer shows and the version block Properties shows
(CompanyName Igneum, ProductName Igneum Miner, FileDescription per exe, 0.3.0, LegalCopyright Igneum contributors),
like the Mac app and DMG already do.
- igneumd.exe, igneum-miner.exe: packaging/windows/embed-resources.sh now takes the worktree and target dir
(defaults vendor/igneum-node-v4 and vendor/igneum-node/target-integration) and relinks with a linker shim first in
PATH instead of `cargo rustc -- -C link-arg`: cargo rustc takes one package and unifies features differently from
the two-package cross-build (300 crates differ), and a configured linker is fingerprinted and rebuilds everything;
the PATH shim changes neither. .rc files moved to 0.3.0.
- igneum-worker-cuda.exe, igneum-worker-opencl.exe: proto-cuda/nvrtc/build-windows.sh compiles the two new .rc
files with windres and links the objects; the icon is made with make-icons.py if missing.
- igneum-app.exe: app/igneum-app/build.rs runs windres on resources/igneum-app.rc for Windows targets and links it
(cargo:rustc-link-arg-bins, no crate dependency); it fails the build if the .rc version drifts from Cargo.toml.
- packaging/windows/resources/verify-exe.py: small PE parser that checks the .rsrc section, icon group and version
strings on the Mac; both build scripts call it, and it checks every exe inside the packages.
- proto-opencl/cl_dynamic.h: clGetEventInfo added to the run-time loader table (4714b51 added the call in host.c
without it, so the one-click OpenCL worker no longer linked under IGNEUM_CL_DYNAMIC).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
After PC 2's gfx1036 (prebuilt-generic path) completed 2,000 jobs a second with no hash from 600 s on: every OpenCL
call in host.c's job path is now fatal on error (exit 3, the miner restarts the worker), the dispatch event must read
CL_COMPLETE, a chunk 20x faster per nonce than the running mean or an output buffer unchanged since the previous
dispatch is a fault, and a stats line every 200 jobs carries the live event and buffer counts (a leak over 16 events
or 12 buffers is fatal too). IGNEUM_FAULT_TEST=N exercises the detectors on a healthy device (verified on Apple
OpenCL: the stale-output guard fires on the chunk after the injected fault). The launcher shows 'worker fault' and
'restarting' for the card on the miner's WORKER FAULT line and drops the last rate. The emulation's NVRTC stand-in
now applies NVRTC's execution-space rule (program.h(46) igneum_launch_* declarations are host code unless
-default-device or -DIGNEUM_NO_CUDA), which is what the RTX 5090 reported; test.sh checks the rejection.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
host.c --serve --pack <dir> serves a pack read at run time (packfile.h), whatever pack the exe was built against: the
first pair is built and self-tested like a prepared one (pairBuffers, pairSelfTest), prepared pairs are self-tested
too, the ready line says path prebuilt-generic. cl_dynamic.h (IGNEUM_CL_DYNAMIC) opens OpenCL.dll with LoadLibrary and
fills a function pointer per entry point, so the mingw build links no import library. test-generic.sh checks the mode
through Apple OpenCL with the packs of the CUDA emulation test: PASS (192 found, prepare and swap, 15 sampled hashes
equal igneum-pow hash-bound), built against a different placeholder pack than it served.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a
load's source from the registers written earlier and not read by a load since, the other
48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no
stale load source, every register injected; dynamic: 64 units on the seed-keyed
closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias
within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced
by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the
generator version, attempt and program id. Version 1 kept as generate_v1 for the census.
Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export;
new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of
39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical
programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000);
CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs
at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887.
Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA,
OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
proto-metal --serve compiles igneum_hash_bound at runtime and mines jobs from stdin (init words in buffer 3, program
and dataset cached per seed); proto-cuda and proto-opencl host serve modes from the pack's kernel_bound.cu / .cl with a
seed guard; --vendor device filter for OpenCL. windows-miner/: START-MINING.bat + start-mining.ps1 (GPU and tool
detection, pack export, cached builds, MINERS identities per vendor, status every 30 s, uploads every 60 s, rebuild on
seed change, Ctrl+C summary), README.txt, make-package.sh (igneum-mine-test.zip with a cross-compiled miner).
docs: fork-divergence devnet v1, bench-log entry with the CPU, Metal and overnight numbers.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.
proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.
Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>