Commit graph

21 commits

Author SHA1 Message Date
igneum-labs
141b559b82 Merge origin/master into ca3-coord: Counter ASIC 3.0 complete (every gate green, P2 green, P1 written); the drive-ref check skips single-quoted here-strings and the copied-sources check reads code lines only (master's CI red on dbdfda0); main's decisions and the close in the status file 2026-10-06 18:05:54 +00:00
igneum-labs
e05eb0c43a Counter ASIC 3.0 gates (node): the program id carries the class. A class v4 program is generator 4 wherever it is made: the CLI's --era path stamps the generator from the class (era_generator_of: 4 on V4_CLASS, 3 otherwise; ProgramClass::of_load_class), show honours --program-class and --era-hex; the shadow block marks class v4 in packcheck, packfile.h and the Metal worker (a generator 3 pack with IGNEUM_SHADOW_INSTRS is refused as a v4 program stamped v3, a generator 4 pack without it is refused; generator 2 ladder packs unchanged); the seven gate packs re-exported (generator 4, id c120d7963abdcd96 against the v3 control's 73bcbfe8ccf988f1, every other line byte-identical); class-v4.mjs asserts every v4 epoch's id against the CLI's same-seed v3 and v4 ids (--id-check-against v4 is the assertion's failed case); v2 and v3 ids byte-identical (60 + 7 + 4 + 19 + 7, the pinned packs)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 16:24:02 +00:00
igneum-labs
e925b20f7c Counter ASIC 3.0 gates (node): class v4 = mx8+sh256x27 through the stack (igneum-pow ProgramClass::V4, V4_CLASS, generator 4, program_id(4, seed, attempt), the era composed as v3's; program.h and program.json class v4; packcheck, packfile.h, the CUDA and OpenCL identity rule and the Metal worker accept generator 4 and the class=v4 token, the Metal worker takes v4 from the prepared pack only; v2 and v3 byte-identical, the pinned packs diffed); the G4 harness class-v4.mjs (two switches, the never case); override-60x.json: the v4 field at never, the duplicated proving v1 block removed, the four 0.3.12/0.3.13 fields added; CI check override-json-check.sh (no duplicate key in any override file)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 15:58:49 +00:00
igneum-labs
563463b35f Merge release-0.3.12 (fda4684) into ember-tune: 0.3.11's six-section View and card order kept, Ember Tune's line and switches re-added on it; the tune fields move into hotplug::apply_pref; the power-cap plan keeps present(); both CI test lists; 132 app tests, 26 UI tests, every gate green
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:20:46 +00:00
igneum-labs
39ecd7c52b hot table (Counter ASIC 2.0 layer 5), measured and not adopted, on the ca2-v3 composed class (squash of tag ca2-cache-history-2026-10-05)
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.

Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:57:39 +00:00
igneum-labs
e08909f138 mixer x4 and the cache growth rule (Counter ASIC 2.0, class v3 construction): LoadClass mixer_mult and growth, LoadClass::MX4 (v2 loads, no width roll), memhard::Shape in MixParams, m mixer applications per round with keys round_key(r m + j), Cache::fill_log2, the option C schedule (growth_doublings, cache_log2_words, dataset_log2_words, days_since_genesis) with its test table, day-sized Epoch entries, the three emitters (m loop only for m > 1, v2 text unchanged), program.h and program.json fields, packfile.h mixerMult, packbench and OpenCL host prints, --class mx4 and --days on the CLI; docs/plans/mixer-x4.md design and spec text, docs/analysis/chip-model-v3.md, the 5090 and 9070 XT dataset-build playbooks (measurements and vectors to follow)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:54:10 +00:00
igneum-labs
fe852a4335 AMD telemetry: igneum-gpu-telemetry (ADLX on Windows, amdgpu sysfs on Linux, PDH utilisation fallback) feeds the card row's draw, temperature, fan, memory clock and MH/W; measured on PC 1: 9070 XT 198.9 W, 64 C, 657 rpm, 17.73 MH/s = 0.089 MH/W beside the 5090 at 307.6 W, 122.30 MH/s = 0.398 MH/W
the project lead watched the 9070 XT at 90% usage with its fans barely turning and the app could not say what it drew: the
draw, temperature and MH per watt line came from nvidia-smi only, and the earlier per-watt figure used the board
rating. proto-opencl/gpu-telemetry.c prints one line per AMD card per sample (bus from SetupAPI by the display
device's name, kind, name, watts, temp_c, fan_rpm, fan_pct, mclk_mhz, gclk_mhz, util_pct, source), built by
build-windows.sh against vendor/adlx (the SDK clone), shipped by make-payload.sh and push-inputs.sh. The engine
runs it with -l 5 beside nvidia-smi (Source::AmdTelemetry, tick_amd_telemetry), parse_amd_telemetry fills
power_w, temp_gpu, fan_pct, fan_rpm, mclk_mhz, util_pct and telemetry_at on the AMD card matched by kind and
ordinal, so eff_mhw and the dashboard's existing line show it; app.js shows fan and memory clock when present.
Tests: three on the parser with lines captured on PC 1 and the Mac fixture; the sysfs path ran on a fixture tree.

Measured over 20:27:45 to 20:29:41 UTC with both cards mining (docs/bench-log.md, under the 9070 XT ceiling table):
9070 XT 198.9 W (193 to 212), 64 C, 657 rpm, 2,505 MHz memory, 3,290 MHz shader, 100% busy, 17.73 MH/s =
0.089 MH/W; RTX 5090 307.6 W, 69 C, 44% fan, 122.30 MH/s = 0.398 MH/W.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:43:21 +00:00
igneum-labs
17641afbe1 Workers: the program class and era on the serve protocol; a pack of the wrong class or era is refused (Counter ASIC 2.0)
packfile.h: pf_load refuses a generator other than 2 or 3 (spec 01 section 1.4.5), reads IGNEUM_PROGRAM_CLASS
(must match the generator) and IGNEUM_ERA_SEED_HEX; pf_pack_class_ok and pf_class_token are the one rule for
the `class=<v2|v3> era=<hex>` tokens a job or prepare line of a class v3 epoch ends with (a v2 line is the
line of before, byte for byte). CUDA worker.cpp and OpenCL host.c: a pair's identity includes its class and era
when the line names them, so a prepared pack of the right class wins over the resident pack of the same seeds;
right seeds with the wrong class answer `need` plus `error <id> pack <dir>: program class mismatch ...`, and a
prepare on a pack of the wrong class fails in plain words. Metal main.swift: refuses every class v3 line until
the Swift generator carries version 3 (the integration branch). emu/packfile-test.c: the generator rule, a v3
pack with its era, the token matcher, the contradicting class line (13 checks pass).
igneum-pow verify.rs: Epoch::chain_dataset(day, class), the one entry the node's day cache goes through, so the
ca2-mixer class-specific item construction and cache schedule have a seam.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:33:59 +00:00
igneum-labs
d28a32e403 Merge commit '4163143' into ca2-v3
# Conflicts:
#	proto-cuda/nvrtc/packfile.h
2026-10-05 20:17:46 +00:00
igneum-labs
a094972548 read-width: packfile accepts a string-seed pack (any even-length hex epoch seed; the seed-word re-derivation still checks it); bench-only playbooks for the second PC round
Round 1 (run-readwidth-{5090,9070}-20261005) delivered the probes and refused every pack: pf_load demanded the chain's 32-byte epoch seed and the experiment packs carry igneum-pow --seed strings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:09:58 +00:00
igneum-labs
8c898d3f60 read-width: CUDA worker --bench and --memprobe, scratch arena from the occupancy capacity; PC playbooks (card under test off in the app, restored after); harness arena label
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:05:57 +00:00
igneum-labs
22dc44141d read-width: scratch per warp is a class parameter (32 or 128 KiB, under the 6 GB working-set cap), distinct-address rule bounds dataset loads only; Metal pack harness; OpenCL --bench-pack, scratch args and 16-byte probe; packfile class fields; OpenCL emulator persistent launch
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:58:20 +00:00
igneum-labs
41631431c0 Workers: a pack's seed words are its program attempt's words, not the bare seed's (epoch 34 incident, 5 October 2026)
From 18:23Z both Windows workers (CUDA on PC 1 and PC 2, OpenCL on PC 1 after the 18:34Z node restart)
refused every pack for epoch 34 with "the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT", and
the miner and the app restarted them every 5 to 60 s until 18:44Z and beyond. The packs were correct.
The generator retries a rejected candidate with seed || k_le32 (attempt_words); epoch 34's attempt 0 was
rejected (246 of 16384 final register values saturated, limit 163) and attempt 1 accepted, so the pack
carried attempt 1's words while pf_load (proto-cuda/nvrtc/packfile.h) derived the expected words from
the bare seed. Both workers also matched jobs to pairs by those bare-seed words, so even a loaded pack
of a retried program would have answered "epoch seed mismatch" on every job.

The rule, in one place per language:
- packfile.h: pf_program_words(bytes, attempt); pf_load reads IGNEUM_PROGRAM_ATTEMPT and checks the
  attempt's words; the refusal says "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not
  attempt N of the epoch seed ..." in plain words.
- worker.cpp and proto-opencl/host.c: a job belongs to a pair when the seed hex the node sent is the
  pair's (pairIs); the compiled-in placeholder pack keeps the word comparison.
- igneum-pow/src/packcheck.rs: verify_pack_texts / verify_pack_dir, the same rule in Rust; the miner
  checks every pack it writes with it before a worker sees it (vendor/igneum-node pack-loop branch).
  Tests pin the attempt vectors of epoch 34 on both sides (one vector, two implementations), that
  epoch 34 is attempt 1 and epoch 33 attempt 0, a known-good pack of a later attempt, a known-mismatched
  (out of date) pack, and self-contradicting packs.
- proto-cuda/nvrtc/emu/packfile-test.c (+ .sh, in CI): pf_load on a known-good attempt-1 pack, the
  checked-in attempt-0 pack, and the known-mismatched bare-words pack.

The app (app/igneum-app):
- watchdog.rs: PACK_OUT_OF_DATE_CODE 44, PackRebuilds (at most 3 pack exports per epoch, then the card
  shows the reason), pack_refusal (the worker's "error 0 pack" line and the miner's "PACK OUT OF DATE"
  line), pack_epoch_of; tests on the incident lines, known-good and known-mismatched.
- engine.rs: exit 44 exports the pack again before the restart instead of a blind restart, the strip
  says "program pack out of date, rebuilding", the card and the log name the condition; at the cap the
  card is marked failed with the reason and tries again in 10 minutes.

The relay (relay/lib/parse.mjs): PACK_MISMATCH; the card reads "pack mismatch, rebuilding (N refusals
in the tail, M restarts)" in `node tools/console.mjs machines` instead of a bare restart count; tests
on the PC 2 tail of 18:27Z and a healthy tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:10:07 +00:00
igneum-labs
382039ea5d Variant race: bench-log entry (Metal on the M5 Max, 14 variants, two programs, serve check), the base-only short-circuit, the PC 1 zip's hash in the plan
Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again
(9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow).
Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two
programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 123ee1a).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:37:27 +00:00
igneum-labs
123ee1a2ee Variant race: a 150 ms pause between timed windows so the job loop gets the card (the mutex is not fair; a queued job waited the whole race, 36 s, in the first serve check)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:18:04 +00:00
igneum-labs
27bcb81494 Variant racing: the NVRTC and Metal workers compile several kernel variants at every prepare and keep the fastest for the hour
the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline
built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs
on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple;
2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's
vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two
seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped,
always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so
far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time.
A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model.

NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the
pack format, the miner and igneum-pow are untouched and old packs race. --race --pack <dir> runs the race alone.
Under IGNEUM_EMU the race is off (the stand-in checks the handed-over text is the pack's). Metal: the MSL hooks in
generateMSL, a lock-protected kernel slot per program, a pair compiled inline races after its first job,
--race-test --seed --day runs the race alone with a table. OpenCL is not raced yet. docs/design/miner-tuning.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:08:45 +00:00
igneum-labs
4d907d8a14 Workers answer a job they have no pair for with a need line; the miner prepares the current pair (PC 2 stuck on the previous epoch)
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 13:32:58 +00:00
igneum-labs
d69d70c3fc Windows exes carry the coin icon and a version block (the project lead's rule, 4 Oct 2026)
Every Windows executable now ships with the coin icon Explorer shows and the version block Properties shows
(CompanyName Igneum, ProductName Igneum Miner, FileDescription per exe, 0.3.0, LegalCopyright Igneum contributors),
like the Mac app and DMG already do.

- igneumd.exe, igneum-miner.exe: packaging/windows/embed-resources.sh now takes the worktree and target dir
  (defaults vendor/igneum-node-v4 and vendor/igneum-node/target-integration) and relinks with a linker shim first in
  PATH instead of `cargo rustc -- -C link-arg`: cargo rustc takes one package and unifies features differently from
  the two-package cross-build (300 crates differ), and a configured linker is fingerprinted and rebuilds everything;
  the PATH shim changes neither. .rc files moved to 0.3.0.
- igneum-worker-cuda.exe, igneum-worker-opencl.exe: proto-cuda/nvrtc/build-windows.sh compiles the two new .rc
  files with windres and links the objects; the icon is made with make-icons.py if missing.
- igneum-app.exe: app/igneum-app/build.rs runs windres on resources/igneum-app.rc for Windows targets and links it
  (cargo:rustc-link-arg-bins, no crate dependency); it fails the build if the .rc version drifts from Cargo.toml.
- packaging/windows/resources/verify-exe.py: small PE parser that checks the .rsrc section, icon group and version
  strings on the Mac; both build scripts call it, and it checks every exe inside the packages.
- proto-opencl/cl_dynamic.h: clGetEventInfo added to the run-time loader table (4714b51 added the call in host.c
  without it, so the one-click OpenCL worker no longer linked under IGNEUM_CL_DYNAMIC).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:44:20 +00:00
igneum-labs
4714b510f2 OpenCL worker fault guard, miner-side fault state in the launcher, NVRTC annotation rule in the emulation
After PC 2's gfx1036 (prebuilt-generic path) completed 2,000 jobs a second with no hash from 600 s on: every OpenCL
call in host.c's job path is now fatal on error (exit 3, the miner restarts the worker), the dispatch event must read
CL_COMPLETE, a chunk 20x faster per nonce than the running mean or an output buffer unchanged since the previous
dispatch is a fault, and a stats line every 200 jobs carries the live event and buffer counts (a leak over 16 events
or 12 buffers is fatal too). IGNEUM_FAULT_TEST=N exercises the detectors on a healthy device (verified on Apple
OpenCL: the stale-output guard fires on the chunk after the injected fault). The launcher shows 'worker fault' and
'restarting' for the card on the miner's WORKER FAULT line and drops the last rate. The emulation's NVRTC stand-in
now applies NVRTC's execution-space rule (program.h(46) igneum_launch_* declarations are host code unless
-default-device or -DIGNEUM_NO_CUDA), which is what the RTX 5090 reported; test.sh checks the rejection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:42:36 +00:00
igneum-labs
734dbd134f NVRTC worker: -default-device (NVRTC rejects unannotated pack helpers as host code; found on the first real card)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:12:35 +00:00
igneum-labs
f535abb59d One-click CUDA worker: driver API + NVRTC loaded at run time, pack read at run time, Mac emulation check
proto-cuda/nvrtc/worker.cpp serves igneum-miner's worker protocol with nothing installed but the NVIDIA driver:
nvcuda.dll and the redistributable nvrtc64_120_0.dll are opened with LoadLibrary (cuda_api.h), the pack's kernel.cu
and kernel_bound.cu are handed to NVRTC byte for byte up to the host launch wrappers with program.h and memhard.h as
named headers, every pack is self-tested against its vectors.h (cache head, last line, FNV-1a 64, dataset head, last
word, 64 samples, 96 vector lanes) before it serves a job, prepare runs on a thread for the hourly swap, and a job on
seeds without a pair makes the worker find the miner's pack by seeds.txt and build it. packfile.h (C99) reads a pack
directory. fetch-redist.sh verifies and stages the NVIDIA 12.8.93 redistributables and the Khronos headers
(THIRD-PARTY.md records URLs, hashes and the EULA clause); build-windows.sh cross-compiles with mingw (static,
KERNEL32 + Universal CRT only). emu/: the two libraries as host functions, the pack kernels on host threads, the NVRTC
source compared with the pack files; test.sh PASS on the Mac (two real packs, swap, self-heal, 17 sampled hashes equal
igneum-pow hash-bound). Not run here: the real NVRTC compile and driver load (RTX 5090 PC).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 09:55:05 +00:00