igneum/proto-cuda
igneum-labs 7eed16a29a Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so
The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh).

The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge.

Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:39:50 +00:00
..
emu era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook 2026-10-05 21:41:33 +00:00
nvrtc Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so 2026-10-07 18:39:50 +00:00
packs Lottery hash: generator version 2 (16 load slots, fresh sources, acceptance rule), every vector re-cut, packs regenerated, three workers re-checked, 20,000-program census 2026-10-04 07:52:40 +00:00
packs-ca2-era era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook 2026-10-05 21:41:33 +00:00
packs-ca2-hot hot table (Counter ASIC 2.0 layer 5), measured and not adopted, on the ca2-v3 composed class (squash of tag ca2-cache-history-2026-10-05) 2026-10-05 21:57:39 +00:00
packs-ca2-mixer verifier regression of e08909f fixed (derive_items out of line, one instance per cache size with the line mask a constant: v2 0.609 ms per unit against readwidth's 0.607, was 1.33); x8 into class v3 under the delegated rule (V3_CLASS = MX8; mx8-genesis and mx8-devnet-epoch0 re-exported through the seam, the devnet one with the era inside; the mx4 packs kept as the x4 record, generator 2); tests/packs.rs and tests/mixer.rs on the x8 class; mixer-x4.md 6.2a, 6.4a, 6.5 decision, 6.6 the regression; chip-model-v3.md headline x8 (0.92x with the factor); bench-log addendum with the PC 1 build rows 2026-10-05 22:10:55 +00:00
packs-ca3-derive Counter ASIC 3.0 item 2: the per-day item-derivation program (class dr736), its interpreter, emitter and packs 2026-10-06 07:41:23 +00:00
packs-ca3-shadow Counter ASIC 3.0 gates (hash): class v4 sub-version 2 (AP-F8-1, main's ruling B2: 0.3.20 ships sub-version 1 untouched; this stream is object byte 7). F1, the draw: a load's source is drawn only from registers fresh by dataflow (fresh at the start; a load keeps freshness only from a fresh source; add, sub, xor, mad, shfl from either operand; rotl, rotr from their operand; or, mul, mulhi never), keyed on the class v4 shape on EVERY draw path (era or not, the pass count set aside), so a census through candidate_class reads the chain's stream. (a'), accept.rs: the same freshness run to its fixpoint over the loop (base then shadow block) and every load's source fresh in the steady state, else the candidate is rejected and the next attempt drawn (closes the iteration boundary the draw cannot see: F8's p11, an or at 63 feeding a load at 1, and the load-after-load and rotate-of-saturated chains of p6, p23, p26, p31, p34). (c'), accept.rs: per load site, the count of source values equal to 0 or all-ones over the 64 units' 16,384 evaluations, rejected at 164 or more (the (c) limit), the backstop for any delivery of saturation (zero and all-ones alike: p45's mulhi zero). Both keyed on the class v4 shape, so v2 and v3 verdicts and ids do not move. PROGRAM_SUBVERSION_V4 = 2 in the id suffix and the pack lines. The seven gate packs re-exported: the devnet epoch-0 seed's attempt 0 is now rejected and attempt 1 accepted, id a788661687db4bb3 (must-differ: c120d7963abdcd96 the 6 October stream, 1a4230699a6b9c60 sub-version 1); the seven 256-block ladder packs of packs-ca3-shadow re-exported under the rule (the no-era path moves too; their measured rates stand as the old stream's). Tests: the generator test checks the fixpoint rule on the amended program and the known-failed case (the v3 stream re-labelled v4); the mixer contract checks the dataflow rule on every V4-shaped class, era or not 2026-10-07 13:01:40 +00:00
packs-ca3-v4 Counter ASIC 3.0 gates (hash): class v4 sub-version 3, first commit (AP-F8-3; main's word: 0.3.21 ships byte 5, sub-version 3 is 0.3.22's). The acceptance executes the latency-shadow block as the hash does: run_unit runs the block after instruction 63 of every iteration, reps times with the iteration's sel (verify.rs); until now it ran the 64 base instructions only, so every dynamic acceptance test on a class v4 program judged a program the chain never hashes, which is the whole residual class behind sub-version 2's 8 of 64 gate failures. The test acceptance_executes_the_shadow_block_as_the_verifier_does pins the acceptance's execution to verify.rs on the devnet epoch-0 program and the six test eras (equal output bit counts over the 64 units; different with the shadow stripped). PROGRAM_SUBVERSION_V4 = 3 (a new verdict is a new stream). The seven gate packs re-exported: the devnet epoch-0 seed still accepts at attempt 1, so its program and fingerprint are sub-version 2's (e370fb2080b7dbb1) under the new id a785001687d8688a (must-differ: c120d7963abdcd96, 1a4230699a6b9c60, a788661687db4bb3); the seven 256-block ladder packs re-exported. The ledger entry carries AP-F8-2's exhaustion half as FIXED-AND-PASSED at fbb00320 (0 of 10^6, max attempt 35, r = 0.67), AP-F8-3, and p23's localisation (site 7 reads r6 = (mulhi | r4) ^ r4 = r6 & ~r4; 0.84 of uniform distinct indices at 2^20, 0.55 at 2^24, reproduced in the acceptance's own execution); the second commit, a per-site distinct-index ratio, is held 2026-10-07 14:20:02 +00:00
packs-readwidth era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook 2026-10-05 21:41:33 +00:00
windows-app Rotation phase 2: packagers read the intake key and the downloads token from files (IGNEUM_INTAKE_KEY_FILE, IGNEUM_DL_TOKEN_FILE, .next by default), no key literal in the tree, app header line with fingerprints, ship-app --dl-both, logs --rotation, tools/repo/fresh-repo.sh with the dry run, docs/plans/rotation-phase-2.md 2026-10-05 07:43:44 +00:00
windows-miner Rotation phase 2: packagers read the intake key and the downloads token from files (IGNEUM_INTAKE_KEY_FILE, IGNEUM_DL_TOKEN_FILE, .next by default), no key literal in the tree, app header line with fingerprints, ship-app --dl-both, logs --rotation, tools/repo/fresh-repo.sh with the dry run, docs/plans/rotation-phase-2.md 2026-10-05 07:43:44 +00:00
windows-node Reproducible builds: SOURCE_DATE_EPOCH from the commit's author time, TZ=UTC and one fixed target path in every build path; self-test 2026-10-06 20:09:23 +00:00
.gitignore Igneum: design docs, Metal lottery-hash prototype, CUDA test pack, finality simulation 2026-10-03 15:06:01 +00:00
build.bat windows-miner: CUDA build finds the MSVC and SDK libraries itself, uploader reads logs in use, deeper toolset search 2026-10-03 19:38:22 +00:00
build.sh GPU workers on the real hash: serve protocol in Metal, CUDA and OpenCL, Windows mining package, devnet v1 log 2026-10-03 19:18:19 +00:00
CHECKLIST.md Memory-hard dataset: 256 MiB ChaCha cache, 8 dependent reads per item, CPU verifier on the cache, levers, CUDA pack igneum-genesis-mh 2026-10-03 16:12:00 +00:00
dot4-probe.cu Counter ASIC 2.0 layer 6: SRAM mirror analysis (cited bit cells N7 to N2 and 18A, area and cost per node, no cache growth rule in the spec, options A to E for the project lead); layer 7 dp4a probes for Metal, CUDA and OpenCL (standalone, no lottery kernel) 2026-10-05 20:09:10 +00:00
family-probe.cu Counter ASIC 3.0 items 6 and 7: family step-cost probes (Metal, CUDA), the PC 2 playbook, the M5 Max rows in the bench-log 2026-10-06 07:36:12 +00:00
host.cu era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook 2026-10-05 21:41:33 +00:00
README.md Windows package 0.3.0: prebuilt one-click workers first, nvcc/cl.exe build path as the fallback 2026-10-04 09:55:05 +00:00
WINDOWS-MINER.md Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so 2026-10-07 18:39:50 +00:00

igneum-bench-cuda (proto-cuda)

The NVIDIA twin of proto-metal. It runs the same random-program proof-of-work kernels on a CUDA GPU and checks them bit for bit against results produced on the Mac.

This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the GPU against known answers, and times the kernel. Nothing here earns anything.

Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 vectors PASS each, see docs/bench-log.md). The memory-hard pack igneum-genesis-mh (added later the same day, construction in ../proto-metal/MEMHARD.md) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no NVIDIA figure for the memory-hard dataset exists yet.

Status on 4 October 2026, generator version 2: every pack was regenerated by igneum-pow export (the Rust crate is now the pack source; the Swift exporter is a cross-check) with the adopted generator rule (exactly 16 loads per program, fresh-source loads, the acceptance rule of spec 01 section 1.4.6). All four packs pass the clang emulation (emu/emu.sh, 96/96 standalone, 2 warps per block in batch) and Apple OpenCL on the M5 Max; Metal cross-checked the exported genesis pack and a 2,000-program fuzz. The version 1 vectors, including the 5090's 192 of 192 from 3 October, are retired: the kernel text is unchanged, so those runs remain evidence that the ops agree on NVIDIA, but the first 5090 run on a version 2 pack is still owed (docs/bench-log.md, 4 October 2026 entry). program.h now carries IGNEUM_GENERATOR 2, IGNEUM_PROGRAM_ATTEMPT and IGNEUM_PROGRAM_ID; a worker built from a version 1 pack cannot serve a version 2 node.

The one-click worker (nvrtc/, 4 October 2026)

nvrtc/worker.cpp is the worker the Windows package ships as igneum-worker-cuda.exe: nothing to install but the NVIDIA driver. It opens nvcuda.dll (the driver API, in every driver) and nvrtc64_120_0.dll (NVIDIA's runtime compiler, redistributable, shipped next to the exe with nvrtc-builtins64_128.dll; nvrtc/THIRD-PARTY.md) with LoadLibrary/GetProcAddress (nvrtc/cuda_api.h), reads a pack directory at run time (nvrtc/packfile.h: program.h, seeds.txt, vectors.h), hands NVRTC the pack's kernel.cu and kernel_bound.cu up to the host launch wrappers with program.h and memhard.h as named headers, byte for byte, loads the cubin through the driver, fills the cache, builds the dataset and self-tests both plus the three vector warps against vectors.h before it serves a job. It speaks the same --serve protocol as host.cu (jobs, prepare in the background for the hourly swap, found/done); a job on seeds it has no pair for makes it look for the miner's pack by seeds.txt and build it in the foreground. host.cu stays the ahead-of-time harness and the launcher's fallback when the prebuilt worker is missing and a toolkit exists.

nvrtc/
  worker.cpp         the worker (C++17, driver API + NVRTC loaded at run time, cross-compiled with mingw)
  cuda_api.h         the function-pointer table for the two libraries
  packfile.h         C99 reader for a pack directory: sizes, seeds, vectors; the self-test verdict (shared with proto-opencl/host.c --pack)
  fetch-redist.sh    downloads and verifies cuda.h, nvrtc.h, the two NVRTC DLLs and the Khronos OpenCL headers into nvrtc/redist/ (gitignored)
  build-windows.sh   cross-compiles igneum-worker-cuda.exe and proto-opencl/igneum-worker-opencl.exe (static, KERNEL32 + Universal CRT only)
  THIRD-PARTY.md     every third-party file, its URL, version, SHA-256 and licence clause
  emu/               the Mac check: emu_backend.cpp (driver API and NVRTC as host functions, the pack's kernels on host threads,
                     the NVRTC source compared byte for byte with the pack), test.sh, serve-check.sh (shared with proto-opencl/test-generic.sh)

What was proven on the Mac (4 October 2026): emu/test.sh builds the worker against two real packs (the devnet epoch 0 pack and a second epoch and day exported by igneum-pow), runs --check (self-test PASS, 96 of 96 lanes) and a --serve session: 64 + 64 + 32 found on pack A including the nonces either side of the 32-bit boundary, pack B prepared in the background with its self-test PASS, the swap, the self-heal rebuild of pack A, 17 sampled hashes equal to igneum-pow hash-bound, and the NVRTC source check PASS for all four files (kernel.cu 7,385 of 8,893 bytes with the 1,508-byte host wrapper tail dropped, kernel_bound.cu 6,003 of 6,887, program.h and memhard.h identical). The exe cross-compiles and links with mingw. Not proven here: that the real NVRTC accepts the text with the two stub headers and that the driver runs the cubin; that is the RTX 5090 run (windows-app/TEST.md).

Layout

proto-cuda/
  host.cu              host program: device info, dataset fill, self-test, vectors, bench, size sweep
  build.sh             Linux build (nvcc)
  build.bat            Windows build (nvcc + Visual Studio Build Tools)
  CHECKLIST.md         Metal/CUDA equivalence, op by op, and what was verified where
  packs/<seed>/        one program pack per seed, written by igneum-pow export (generator v2, 4 October 2026)
    kernel.cu          the program as a CUDA kernel, plus fill kernel and host launch wrappers
    kernel.cl          the same program as OpenCL C for proto-opencl (AMD and any other OpenCL device), built at runtime
    program.h          seed, day words, dataset size, loads per hash, wrapper declarations (C99-safe: proto-opencl/host.c includes it too)
    vectors.h          expected outputs for 3 warps (96 x 64-bit) and dataset self-test values
    program.json       the instruction list and all constants, for any other implementation
    vectors.json       the same vectors as JSON
    program.metal      the Metal source the Mac ran, for diffing by eye
    memhard.h          memory-hard packs only: the cache fill and item derivation core, compiled for device and host
    memhard.metal      memory-hard packs only: the Metal cache-fill and build kernels the Mac ran
  emu/                 CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
  nvrtc/               the one-click worker (driver API + NVRTC at run time), its Mac emulation and the third-party record

Four packs are checked in, all generator version 2 with 128 loads per hash. igneum-genesis and igneum-hourly use the original closed-form dataset (IGNEUM_DATASET_MODE 0, implied when the macro is absent). igneum-genesis-mh is the same program as igneum-genesis over the memory-hard dataset (IGNEUM_DATASET_MODE 1): a 256 MiB cache of chained ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache reads through a seed-parameterised mixer (../proto-metal/MEMHARD.md). The hash kernel text is identical in both packs; only the dataset contents differ, so the 96 expected outputs differ. All three were cross-checked on the Mac's Metal GPU before being written.

Prerequisites

Linux

  • An NVIDIA driver recent enough for the toolkit. For CUDA 12.8 that is the R570 series or newer (approximate, from memory; nvidia-smi prints the driver's maximum supported CUDA version in its header).
  • CUDA Toolkit 12.8 or newer. Blackwell (sm_120, RTX 50 series) is not known to older toolkits.
  • A host compiler the toolkit supports (gcc 11 to 13 for 12.8, approximate).

Windows

  • The same driver requirement.
  • CUDA Toolkit 12.8 or newer. Tick the Visual Studio integration in the installer.
  • Visual Studio 2022 Build Tools with the "Desktop development with C++" workload. nvcc needs cl.exe.
  • Run build.bat from an "x64 Native Tools Command Prompt for VS 2022" so cl.exe is on PATH.

Build

Linux:

cd proto-cuda
./build.sh                         # pack igneum-genesis, -arch=sm_120
./build.sh igneum-hourly           # the second pack
./build.sh igneum-genesis-mh       # the memory-hard pack (needs 256 MiB more device memory for the cache)
./build.sh igneum-genesis native   # if sm_120 is refused, let nvcc pick the installed GPU

Windows (x64 Native Tools Command Prompt):

cd proto-cuda
build.bat
build.bat igneum-hourly
build.bat igneum-genesis-mh
build.bat igneum-genesis native

Both scripts run this one command (paths adjusted for the pack):

nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu

Notes

  • -arch=sm_120 is Blackwell. -arch=native (CUDA 11.6 or newer) compiles for whatever GPU is in the machine and is the fallback if the toolkit is too old to know sm_120 (which means it is too old for a 5090 anyway: upgrade).
  • If nvcc on Windows refuses the Visual Studio version, add -allow-unsupported-compiler to the nvcc line.
  • The host code is plain C++17 and the CUDA runtime API. No NVRTC, no third-party libraries, no JSON parser. The kernel is compiled ahead of time from the pack.

Run

./igneum-bench-cuda-igneum-genesis                     # 1 GiB dataset, 5 batches x 2^24 nonces, vectors checked
./igneum-bench-cuda-igneum-genesis --sweep             # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
./igneum-bench-cuda-igneum-genesis --block-warps 4     # 4 warps per block instead of 1 (still bit-exact)
./igneum-bench-cuda-igneum-hourly                      # the second program
./igneum-bench-cuda-igneum-genesis-mh                  # memory-hard dataset: cache fill + build, cache check, vectors, bench
./igneum-bench-cuda-igneum-genesis-mh --sweep          # the same sweep over the memory-hard dataset

On Windows the binaries are igneum-bench-cuda-igneum-genesis.exe and so on.

Flags: --dataset-mib N (power of two, default 1024), --sweep, --batch-log2 B (default 24), --batches N (default 5), --block-warps W (default 1, mirrors the Metal run's one SIMD group per threadgroup), --device D.

What it prints, in order:

  1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and resident warps per SM for the kernel, and which dataset construction the pack uses.
  2. Memory-hard packs only: cache fill time on the GPU (twice), cache fill time on the host (one thread, the same memhard.h text), then the cache check: every one of the 2^26 words GPU versus host, the host FNV-1a 64 against the Mac's, and the head and last line against the Mac's. A FAIL here stops nothing but fails OVERALL.
  3. Dataset fill time (closed form, twice, with write GB/s) or dataset build time (memory-hard, twice, with items/s and cache-line reads/s).
  4. Dataset self-test: 16 head words and word [MASK] against values from the Mac, 64 random words against the host formula (closed form) or the host derivation from the host cache (memory-hard), and the Mac's 64 sampled words (memory-hard packs; those inside the current dataset size).
  5. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the first differing lane printed on FAIL.
  6. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s, GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
  7. A summary table in Markdown and OVERALL: PASS or FAIL. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.

Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.

What PASS means

  • Closed-form packs: the CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not every word).
  • Memory-hard pack: the CUDA cache-fill kernel produced, word for word, the same 256 MiB cache as the host and as the Mac (FNV-1a 64, head, last line), and the CUDA build kernel produced the same dataset words as the host derivation and the Mac's samples (sampled, not every word). The whole chain from day key to dataset word agrees across three compilers (Apple Metal, host C++, NVIDIA CUDA).
  • For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
  • With --block-warps W the in-batch check passing shows that packing W warps per block changed nothing.

A FAIL with a small number of differing lanes points at a shuffle; a FAIL in every lane points at an arithmetic op or the dataset. Send the whole printout either way.

Sending results back

Copy the printed header lines (GPU, CUDA versions, kernel line, program line) and the summary table into docs/bench-log.md under a dated heading, together with the output of nvcc --version and the driver version from nvidia-smi. Keep the full stdout as well. Run both packs and the sweep so the log has the same shape as the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and log both.

Checking a pack without a GPU

emu/emu.sh <pack> [flags] compiles host.cu and the pack's kernel.cu as plain C++17 against a shim cuda_runtime.h and runs the kernels on host threads (32 per warp, a barrier inside __shfl_xor_sync). Use small batches (--batch-log2 13 --batches 1). Only PASS/FAIL matters; the rates it prints are noise. This is how the CUDA text was checked on the Mac on 3 October 2026 (all three packs PASS, see CHECKLIST.md). The memory-hard pack builds the 1 GiB dataset on host threads, which takes about a second on the M5 Max; the cache fill on one host thread took 161 ms. It is not an nvcc build and says nothing about NVIDIA hardware.

Regenerating a pack

Since 4 October 2026 the packs are written by the Rust crate, which is the normative implementation:

cd igneum-pow && cargo build --release
./target/release/igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh     # memory-hard (default)
./target/release/igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
./target/release/igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly
./target/release/igneum-pow export --epoch-hex <32-byte epoch seed> --day-hex <day bytes> --out ../proto-cuda/packs/<name>
cargo test --release                                            # the packs must match the emitters byte for byte

The fourth form is the chain's derivation (igneum-devnet-v4-epoch0: the devnet genesis hash and the day bytes of 2026-10-04). The Metal cross-check is a separate step: proto-metal/igneum-bench --seed <seed> --export-pack <scratch dir> derives the same program with the Swift generator, runs the Metal kernel for the three vector warps and refuses to write unless all 96 outputs match its CPU interpreter; diff its program.json instruction list and vectors.json against the Rust pack (identical on every seed tried, docs/bench-log.md). --day and --dataset-log2 change the dataset constants and are recorded in the pack.