Commit graph

5 commits

Author SHA1 Message Date
igneum-labs
56d6dc6efe Windows package 0.3.0: prebuilt one-click workers first, nvcc/cl.exe build path as the fallback
igneum-common.ps1 uses igneum-worker-cuda.exe (with nvrtc64_*_0.dll next to it) and igneum-worker-opencl.exe when they
are in the folder, passes the exported pack with --pack, and only surveys the toolchain (Find-Toolchain) when a worker
is missing or FORCE_BUILD=1; the dashboard and the status block name the path per card; a prebuilt worker that is not
ready after 150 s falls back to the build path once when a toolchain exists; 90 s of seed mismatch errors re-export
the pack and restart the vendor; WORKER_ARCH overrides the NVRTC target. make-package.sh ships the two exes, the two
NVRTC DLLs, the licence texts, THIRD-PARTY.md and TEST.md (what the first RTX 5090 run should print and what to send
back). README.txt, the bats, proto-cuda/README.md (nvrtc/ section), WINDOWS-MINER.md and the bench log updated with
what the Mac measured.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 09:55:05 +00:00
igneum-labs
fdcab858e3 Lottery hash: generator version 2 (16 load slots, fresh sources, acceptance rule), every vector re-cut, packs regenerated, three workers re-checked, 20,000-program census
igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a
load's source from the registers written earlier and not read by a load since, the other
48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no
stale load source, every register injected; dynamic: 64 units on the seed-keyed
closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias
within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced
by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the
generator version, attempt and program id. Version 1 kept as generate_v1 for the census.

Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export;
new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of
39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical
programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000);
CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs
at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887.

Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA,
OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:52:40 +00:00
igneum-labs
660f0eb16c proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.

proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.

Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 16:52:24 +00:00
igneum-labs
3b00e8b480 Memory-hard dataset: 256 MiB ChaCha cache, 8 dependent reads per item, CPU verifier on the cache, levers, CUDA pack igneum-genesis-mh
proto-metal: default dataset is now the memory-hard construction (MEMHARD.md), --closed-form keeps the original.
Cache fill 2 ms GPU / 185 ms one CPU core; dataset build 20.6 ms; GPU cache == CPU cache on all 2^26 words.
Shortcut ratio: inline kernel 111x faster than honest (closed form) to 4.8x slower (memory-hard), 1 GiB.
CPU verify 0.63 to 0.80 ms per warp at 104 loads, 1.21 ms at 144 loads (4,608 items): 10 ms gate met.
Levers --load-weight and --wide-frac implemented and measured, both off; default generator unchanged.
Fuzz 200/200, edge, determinism, memcheck, stats re-run on the new dataset, all PASS.
proto-cuda: host.cu handles both dataset modes; new pack igneum-genesis-mh with memhard.h; clang emulation PASS
including the three-way cache check. Old packs unchanged; closed-form export is byte-identical to them.
docs/bench-log.md: dated summary.

Note: a concurrent session running git commit -a swept earlier states of these files into its site commits
(7b28d5e through d6539fa); this commit carries the remainder.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 16:12:00 +00:00
igneum-labs
ff263245ac Igneum: design docs, Metal lottery-hash prototype, CUDA test pack, finality simulation
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 15:06:01 +00:00