Commit graph

18 commits

Author SHA1 Message Date
igneum-labs
838fed7989 Counter ASIC 3.0 PC 1 AMD: OpenCL family probe (every 1.13.2 family through AMD's builtins, khr/intel sub-groups, media ops, C forms and local emulation; bit-exact on Apple OpenCL) and the clBuildProgram time line in the OpenCL worker
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 15:39:51 +00:00
igneum-labs
19b30e31f0 Merge remote-tracking branch 'origin/master' into ca2-v3
# Conflicts:
#	docs/bench-log.md
#	proto-opencl/host.c
2026-10-05 22:34:07 +00:00
igneum-labs
39ecd7c52b hot table (Counter ASIC 2.0 layer 5), measured and not adopted, on the ca2-v3 composed class (squash of tag ca2-cache-history-2026-10-05)
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.

Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:57:39 +00:00
igneum-labs
1558971521 era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook
Squashed from six commits (tag ca2-era-pre-squash) for one merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:41:33 +00:00
igneum-labs
e08909f138 mixer x4 and the cache growth rule (Counter ASIC 2.0, class v3 construction): LoadClass mixer_mult and growth, LoadClass::MX4 (v2 loads, no width roll), memhard::Shape in MixParams, m mixer applications per round with keys round_key(r m + j), Cache::fill_log2, the option C schedule (growth_doublings, cache_log2_words, dataset_log2_words, days_since_genesis) with its test table, day-sized Epoch entries, the three emitters (m loop only for m > 1, v2 text unchanged), program.h and program.json fields, packfile.h mixerMult, packbench and OpenCL host prints, --class mx4 and --days on the CLI; docs/plans/mixer-x4.md design and spec text, docs/analysis/chip-model-v3.md, the 5090 and 9070 XT dataset-build playbooks (measurements and vectors to follow)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:54:10 +00:00
igneum-labs
17641afbe1 Workers: the program class and era on the serve protocol; a pack of the wrong class or era is refused (Counter ASIC 2.0)
packfile.h: pf_load refuses a generator other than 2 or 3 (spec 01 section 1.4.5), reads IGNEUM_PROGRAM_CLASS
(must match the generator) and IGNEUM_ERA_SEED_HEX; pf_pack_class_ok and pf_class_token are the one rule for
the `class=<v2|v3> era=<hex>` tokens a job or prepare line of a class v3 epoch ends with (a v2 line is the
line of before, byte for byte). CUDA worker.cpp and OpenCL host.c: a pair's identity includes its class and era
when the line names them, so a prepared pack of the right class wins over the resident pack of the same seeds;
right seeds with the wrong class answer `need` plus `error <id> pack <dir>: program class mismatch ...`, and a
prepare on a pack of the wrong class fails in plain words. Metal main.swift: refuses every class v3 line until
the Swift generator carries version 3 (the integration branch). emu/packfile-test.c: the generator rule, a v3
pack with its era, the token matcher, the contradicting class line (13 checks pass).
igneum-pow verify.rs: Epoch::chain_dataset(day, class), the one entry the node's day cache goes through, so the
ca2-mixer class-specific item construction and cache schedule have a seam.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:33:59 +00:00
igneum-labs
d28a32e403 Merge commit '4163143' into ca2-v3
# Conflicts:
#	proto-cuda/nvrtc/packfile.h
2026-10-05 20:17:46 +00:00
igneum-labs
22dc44141d read-width: scratch per warp is a class parameter (32 or 128 KiB, under the 6 GB working-set cap), distinct-address rule bounds dataset loads only; Metal pack harness; OpenCL --bench-pack, scratch args and 16-byte probe; packfile class fields; OpenCL emulator persistent launch
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:58:20 +00:00
igneum-labs
dc84789e34 read-width experiment (gate 1): load classes W=4/16/64, per-load width mix, scratch RMW variant behind a generator flag; 20 packs; CPU verifier and acceptance mirror; emulator shims
Nothing changes for the default class: the pinned packs are byte-identical (tests/packs.rs), the v2 draw stream is untouched.
LoadClass {mix, load_slots, scratch}: fixed widths w16, w64, w64x4 (4 loads of 64 B), era mixes 50/35/15 and 25/50/25 drawn per load with one extra below(100) roll, and the scratch variant scr0/2/4/8 (persistent warps, 1 MiB per warp, tagged lazy fill, measurement only). A wide load reads the W-aligned address and folds every word: x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]. Program ids carry the class. proto-opencl/host.c taken from opencl-rdna4 23810df (--memprobe, select read-back).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:46:54 +00:00
igneum-labs
d5da02a851 Merge pack-loop (4163143) into release-0.3.10: a pack's seed words are its program attempt's words, not the bare seed's (the epoch 34 outage)
# Conflicts:
#	relay/test/parse.test.mjs
2026-10-05 19:23:59 +00:00
igneum-labs
41631431c0 Workers: a pack's seed words are its program attempt's words, not the bare seed's (epoch 34 incident, 5 October 2026)
From 18:23Z both Windows workers (CUDA on PC 1 and PC 2, OpenCL on PC 1 after the 18:34Z node restart)
refused every pack for epoch 34 with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", and
the miner and the app restarted them every 5 to 60 s until 18:44Z and beyond. The packs were correct.
The generator retries a rejected candidate with seed || k_le32 (attempt_words); epoch 34's attempt 0 was
rejected (246 of 16384 final register values saturated, limit 163) and attempt 1 accepted, so the pack
carried attempt 1's words while pf_load (proto-cuda/nvrtc/packfile.h) derived the expected words from
the bare seed. Both workers also matched jobs to pairs by those bare-seed words, so even a loaded pack
of a retried program would have answered "epoch seed mismatch" on every job.

The rule, in one place per language:
- packfile.h: pf_program_words(bytes, attempt); pf_load reads IGNEUM_PROGRAM_ATTEMPT and checks the
  attempt's words; the refusal says "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not
  attempt N of the epoch seed ..." in plain words.
- worker.cpp and proto-opencl/host.c: a job belongs to a pair when the seed hex the node sent is the
  pair's (pairIs); the compiled-in placeholder pack keeps the word comparison.
- igneum-pow/src/packcheck.rs: verify_pack_texts / verify_pack_dir, the same rule in Rust; the miner
  checks every pack it writes with it before a worker sees it (vendor/igneum-node pack-loop branch).
  Tests pin the attempt vectors of epoch 34 on both sides (one vector, two implementations), that
  epoch 34 is attempt 1 and epoch 33 attempt 0, a known-good pack of a later attempt, a known-mismatched
  (out of date) pack, and self-contradicting packs.
- proto-cuda/nvrtc/emu/packfile-test.c (+ .sh, in CI): pf_load on a known-good attempt-1 pack, the
  checked-in attempt-0 pack, and the known-mismatched bare-words pack.

The app (app/igneum-app):
- watchdog.rs: PACK_OUT_OF_DATE_CODE 44, PackRebuilds (at most 3 pack exports per epoch, then the card
  shows the reason), pack_refusal (the worker's "error 0 pack" line and the miner's "PACK OUT OF DATE"
  line), pack_epoch_of; tests on the incident lines, known-good and known-mismatched.
- engine.rs: exit 44 exports the pack again before the restart instead of a blind restart, the strip
  says "program pack out of date, rebuilding", the card and the log name the condition; at the cap the
  card is marked failed with the reason and tries again in 10 minutes.

The relay (relay/lib/parse.mjs): PACK_MISMATCH; the card reads "pack mismatch, rebuilding (N refusals
in the tail, M restarts)" in `node tools/console.mjs machines` instead of a bare restart count; tests
on the PC 2 tail of 18:27Z and a healthy tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:10:07 +00:00
igneum-labs
bd21f2a680 app: one card row per physical GPU across OpenCL platforms, Windows names on the rows, index-free card keys; the OpenCL worker prints the PCI address
PC 1 after Adrenalin 26.9.2 (5 October 2026): two AMD ICDs each listed the iGPU and the RX 9070 XT, so the 0.3.9
app showed five rows for three GPUs (gfx1036, gfx1201, gfx1036, gfx1201 plus the 5090), ran two workers on the one
9070 XT at 9 MH/s each, called the iGPU discrete, and re-enabled it when its key moved from amd:1:gfx1036 to
amd:0:gfx1036.

detect.rs: the Windows path is a pure assemble(Inputs) over nvidia-smi, the worker's --list and the adapter list.
dedupe_platforms keeps one entry per card per vendor: the fuller platform wins (most GPUs, then the newest driver,
then the first listed) and another platform's device survives only at a PCI address the winner lacks (without
addresses: the same code and ordinal); the dropped entries go to the log, never to a worker. resolve_name puts the
Windows adapter name on the row: by PCI bus (the adapter list now carries DEVPKEY_Device_BusNumber and Address),
else by the gfx code's PCI device ids, else by the card names in a gfx table (gfx1201, 1200, 1100, 1101, 1102,
1030, 1031, 1032, 1036, 1035, 1103, 1150, 90c, approximate), else the table's words, else the code. The code stays
in `code`, the key and the row's tooltip; the key is vendor:code with "#2" for a twin, no index. hotplug::pref_for
reads a 0.3.9 key (vendor:index:code) when exactly one matches; the diff matches by key and PCI address, then by
vendor and name or code. The iGPU is integrated and off by default; the choice survives a moved index.

proto-opencl/host.c: --list prints ", pci bb:dd.f" on the info line from CL_DEVICE_TOPOLOGY_AMD or the NVIDIA bus
and slot ids (needs a worker rebuild through proto-cuda/nvrtc/build-windows.sh and push-inputs; an old worker
still works, by code and ordinal).

Tests: detect::tests has the five-row PC 1 list as a fixture (two platforms, the three adapters with DEV_13C0,
DEV_2B85 and DEV_7550) and asserts three rows named NVIDIA GeForce RTX 5090, AMD Radeon(TM) Graphics (integrated,
off) and AMD Radeon RX 9070 XT (discrete, on, 8 identities), the keys, the console line, and the diff against the
one-platform list before the eGPU (iGPU moved, 9070 XT added, nothing removed); plus the list parser with and
without pci, dedupe with and without addresses, name resolution, twins' keys. cargo test -p igneum-app: 96 passed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 18:20:44 +00:00
igneum-labs
4d907d8a14 Workers answer a job they have no pair for with a need line; the miner prepares the current pair (PC 2 stuck on the previous epoch)
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 13:32:58 +00:00
igneum-labs
4714b510f2 OpenCL worker fault guard, miner-side fault state in the launcher, NVRTC annotation rule in the emulation
After PC 2's gfx1036 (prebuilt-generic path) completed 2,000 jobs a second with no hash from 600 s on: every OpenCL
call in host.c's job path is now fatal on error (exit 3, the miner restarts the worker), the dispatch event must read
CL_COMPLETE, a chunk 20x faster per nonce than the running mean or an output buffer unchanged since the previous
dispatch is a fault, and a stats line every 200 jobs carries the live event and buffer counts (a leak over 16 events
or 12 buffers is fatal too). IGNEUM_FAULT_TEST=N exercises the detectors on a healthy device (verified on Apple
OpenCL: the stale-output guard fires on the chunk after the injected fault). The launcher shows 'worker fault' and
'restarting' for the card on the miner's WORKER FAULT line and drops the last rate. The emulation's NVRTC stand-in
now applies NVRTC's execution-space rule (program.h(46) igneum_launch_* declarations are host code unless
-default-device or -DIGNEUM_NO_CUDA), which is what the RTX 5090 reported; test.sh checks the rejection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:42:36 +00:00
igneum-labs
079ed76a2d OpenCL worker: generic --pack serve mode and OpenCL.dll loaded at run time (one-click AMD worker)
host.c --serve --pack <dir> serves a pack read at run time (packfile.h), whatever pack the exe was built against: the
first pair is built and self-tested like a prepared one (pairBuffers, pairSelfTest), prepared pairs are self-tested
too, the ready line says path prebuilt-generic. cl_dynamic.h (IGNEUM_CL_DYNAMIC) opens OpenCL.dll with LoadLibrary and
fills a function pointer per entry point, so the mingw build links no import library. test-generic.sh checks the mode
through Apple OpenCL with the packs of the CUDA emulation test: PASS (192 found, prepare and swap, 15 sampled hashes
equal igneum-pow hash-bound), built against a different placeholder pack than it served.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 09:55:05 +00:00
igneum-labs
fc3e0a2999 Workers: compile-ahead hot swap in the CUDA, OpenCL and Metal hosts, generator v2 port in the Metal host; windows-miner one worker per card; site bench and journey regenerated
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:55:50 +00:00
igneum-labs
27e944d8eb GPU workers on the real hash: serve protocol in Metal, CUDA and OpenCL, Windows mining package, devnet v1 log
proto-metal --serve compiles igneum_hash_bound at runtime and mines jobs from stdin (init words in buffer 3, program
and dataset cached per seed); proto-cuda and proto-opencl host serve modes from the pack's kernel_bound.cu / .cl with a
seed guard; --vendor device filter for OpenCL. windows-miner/: START-MINING.bat + start-mining.ps1 (GPU and tool
detection, pack export, cached builds, MINERS identities per vendor, status every 30 s, uploads every 60 s, rebuild on
seed change, Ctrl+C summary), README.txt, make-package.sh (igneum-mine-test.zip with a cross-compiled miner).
docs: fork-divergence devnet v1, bench-log entry with the CPU, Metal and overnight numbers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 19:18:19 +00:00
igneum-labs
660f0eb16c proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.

proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.

Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 16:52:24 +00:00