The 03:23Z card run compiled the rewritten kernel for the first time and NVRTC refused it: `igneum_hash_bound_unit(d, ou, baseNonc, mas, i, gid)`, the parameter capture's substr length written e - start where the last character's index e needs e - start + 1 (an LF fault, not a CRLF one; CRLF would have missed the anchors). Fixed; the sparse rewrite now strips \r first so LF and CRLF input give the same rewritten bytes. --list-race gains the wrapper's call line checked name by name against the unit function's parameters, and a compile of the whole rewritten text through libnvrtc.so.12 for sm_120 when the library is present (the box's CUDA 12.8; no device). Gate on igneum-build-1 under lease (class measure), 03:55 UTC, the pack as LF and as a CRLF copy: variants=2 names=base,sp43-w32; rewrite=applied bytes=22266 on both; call="igneum_hash_bound_unit(ds, out, baseNonce, mask, iw, gid)" params=6 args=6 names_match=1; nvrtc=libnvrtc.so.12 arch=sm_120 compiled=1 image_bytes=36256 on both; sp11-w32 the same; the failed case is the card run's own line of 03:23Z. mingw exit 0 and clean; exe igneum-worker-cuda-ca4sparse5.exe sha256 a4550202b301faf22f5329c2ab4fa1c0aa6695dbdaf31c974f316dca2524d7d6.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The 02:52Z rerun on ca4sparse3 read "variants 1 base only, no race" on every sparse row: racePair's push looked the pinned name up in the order list through findVariant, whose on-demand sp<N> path answers for any list, so the sparse variant was "found" in an empty order and never pushed. worker.cpp: raceOrder (pure; membership by name) builds the race's order for racePair; --list-race prints, with no device, the order the run's own option handling builds (variants N, names) and, with a pack, whether the named variant's rewrite applies to its bound kernel (bytes, the nonces argument, the unit function), exit 1 when a named variant is missing. emu/variant-test.cpp gains the race-order case; emu/list-race-stubs.cpp links worker.cpp itself as a Linux check binary. Gate on igneum-build-1 under lease (class measure), 03:04 UTC: variant-test PASS (race-order pinned=sp43-w32 variants=2 names=base,sp43-w32; base only 1); the bench's own output `--bench --pack v4-devnet-epoch0 --block-warps 32 --variant sp43-w32 --list-race` reads `variants=2 names=base,sp43-w32` and `sparse_blocks=43 block_warps=32 rewrite=applied bytes=22261 nonces_arg=1 unit_fn=1`, exit 0; the base form reads `race_off=1 variants=1 names=base`; mingw cross-compile exit 0 and clean; exe igneum-worker-cuda-ca4sparse4.exe sha256 84846396559004a8df61881c15ecb42fa3fc1010ad99074e0c0b53e81bb1ca3b.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
worker.cpp: launchGridBlocks (the dispatch's grid: nonces / block, or the sparse block count) and benchRaceOff (a bench keeps the race off only without --variant) as pure functions; --bench with --variant runs the pinned race (base and the named variant, no timing, the variant installed whatever its speed), buildPair races in that case, the bench launches with the served pair's block, the RESULT line carries variant=, sparse_blocks= and block_warps=, and a served kernel other than the requested one prints variant_not_installed. emu/variant-test.cpp: the known-failed pair on the host with no card, against the real class v4 pack: base reads race off and 524,288 blocks of 32; sp43-w32 reads race on, 43 sparse blocks of 32 warps, the rewrite carrying the nonces argument and the unit function, 43 blocks of 1,024 (the 7 October fault read the same shape and "variant base" for both). Gate: igneum-build-1 under lease pool (class measure), 00:46 UTC: variant-test PASS, exit 0; mingw cross-compile exit 0, -Wall -Wextra clean; exe igneum-worker-cuda-ca4sparse3.exe sha256 0ba97edcd5c46a302a7ff5ddd1bbb1e493ca15f64d0757820ed645972df3bb56 (unrun on a card: the hash lane's rerun).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
accept.rs: the per-load class's value-level test (the adv-cache-2 requirement): the one-count of every index bit per load site over the 64 units' 16,384 addresses, BiasedIndexBit outside the 6-sigma band (384 of 8,192); generator: a sub-block writer of the next load's source is never a product (mad joins mul, mulhi, or). The record (tests/ca4_trace.rs, igneum-build-2, 22:44 to 22:52 UTC): over 64 seeds 22 of 1,621 candidates accepted (1.4 percent), 42 of 64 seeds exhaust the chain's 32 attempts; the first failing test: biased index bit 775, duplicate lanes 643, the base rule 110, (b) 43, (a) 28; the genesis seed 0 of 32; candidate 0 of the class carries index bit 0 at 40 of 1,024 at site 0 (z 29.5, sub-block writer rotl of a mad). The class v4 shape across 16 drawn eras: 0 duplicate pairs, with the biased product bit at address bit R exactly in 14 of 17 eras on this pre-amendment generator (the adv-cache-2 class, replicated). The tests record these as reports and bounds; the construction stays in the crate behind the pack as the measured dead end, never a chain class.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The finding (tests/ca4_trace.rs, the first export traced at 22:16 UTC): the per-load class derived 10,728 distinct items of 12,288 over three units (the class v4 shape 12,286), 1,482 same-iteration duplicate lanes at sites 8, 10 and 15; the mechanism is not the sub-block's last writer alone: a lossy base writer (mul, mulhi, or) followed by 27 passes of a 16-instruction map collapses a load's source register to 1 to 17 distinct values in 32 lanes before the next load (the census's 29 failing seeds of 64 at 22:16 UTC name mulhi, mul, or, rotl and load as the last base writers). The fix, two layers: (1) generator: a per-load sub-block instruction that writes the NEXT load's source register is redrawn from the injecting families when its op is mul, mulhi or or (consumes a draw; the class's own stream); (2) accept.rs: the dynamic test steps the per-load sub-blocks inside run_unit (alu_step, the same arithmetic as the arms) so saturation and distinctness are judged on the register file the loads read from, and a new rejection DuplicateLanes (the per-load class only) refuses a candidate whose load reads one address in two lanes of a unit; a rejected candidate redraws the attempt. After: the trace reads 12,287 of 12,288 with 0 duplicate lanes on the genesis seed; the 64-seed census on a second dataset reads 1 duplicate pair in 16,384 rows (seed ca4-census/49, the chance floor of a 2^24 index space, about 0.5 expected; the class v4 shape's own 2 of 12,288 are the same floor). verify.rs gains trace_load_indices (the diagnostic). Suite on igneum-build-2: 64 + 2 + 7 + 4 + 19 + 2 + 7 passed, 0 failed (RESULT rc=0, 22:23 UTC). The pinned packs unchanged; mx8+shl256x27's accepted attempt and id move, so the pack is re-exported.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
igneum-pow: ShadowClass gains per_load (the shadow block split into 16 sub-blocks of instrs / 16, sub-block j run reps times right after the j-th load; name "+shl<S>x<R>", id bytes "perload") and tiles (int8 mma.m8n8k16 u8 tile steps per iteration after the shadow block, [a, b, c, c2] drawn from the stream after the shadow draws, c2 != c, both outputs consumed; name "+mm<R>", id bytes "mm8/" || tiles_le16); ShadowClass::block keeps every class v4 literal and id unchanged. New module mm8: the tile step on the warp register file, the scalar reference and an AVX2 path (maddubs on a 4-way split of B so no 16-bit lane saturates), the unit test pinning SIMD equal to scalar on 64 seeds and the all-ones edge; IGNEUM_MM8_SCALAR=1 forces the reference. verify: the per-load sub-blocks inside the base loop, the tile block after the shadow. emit: the sub-block after each load line in Metal, CUDA and OpenCL; the tile block as inline PTX mma in CUDA and as the shuffle-and-byte-product reference in Metal (mm8_ref) and OpenCL (IGNEUM_MM8_REF with IGNEUM_SHFL_IDX in the three exchange variants); lane declared when tiles are present; program.h gains IGNEUM_SHADOW_PER_LOAD and IGNEUM_MM8_TILES[_PER_HASH]; program.json the placement and the descriptors. Nothing is emitted for a class without the fields: the packs test holds the pinned packs byte for byte and the re-exported mx8+sh256x27 is byte-identical to packs-ca3-shadow/sh256x27 on all seven files.
proto-cuda/packs-ca4: mx8_sh256x27 (control), mx8_shl256x27, mx8_mm128, mx8_mm512, mx8_mm1430 over seed igneum-genesis, day 2026-10-03, 1 GiB, generator 2 (every export OVERALL PASS on igneum-build-2, 21:40 UTC); unrun on a card (the hash lane's PC 1 job).
Gate: igneum-build-2 `cargo test --release -p igneum-pow`: 64 + 7 + 4 + 19 + 2 + 7 passed, 0 failed (RESULT rc=0, 21:42 UTC).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sections 15 to 19: every GPU block gaming and AI paid for, priced against the best public 4 to 5 nm chip figure with its k band, verifier cost and the edge at the measured premium (the one-sentence answer: nothing reads k above 1 with certainty; the int8 tensor tile is the only block near or above 1, so the tensor shadow with a SIMD byte-dot verifier is the new rank 3); the capex column added to the chip model (both sides capex-dominated at 7x to 10x their electricity; the per-unit capex wall unreachable by 3x to 7x; the project wall doubled from USD 100 M to about USD 200 M of break-even cap by the shadow-per-load design, the mission lane's method); the microbench of 20 probes documented as built (section 18); the owed list.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
proto-cuda/nvrtc/worker.cpp, cuda_api.h: `--microbench [--mb-seconds 60] [--mb-only a,b]` runs 20 probes one at a time at full residency for the window, each compiled by NVRTC on its own (a refused form drops that probe with its error on the row): the sleeping-SM floor, int32 add-xor-rotate, multiply, mulhi, byte permute, LOP3, warp shuffle, FP32 FMA, FP16x2 FMA, tensor tiles at u8 m8n8k16, s8 m16n8k32, f16 and bf16 m16n8k16, e4m3 m16n8k32, L2-resident chases at 32 and 64 MiB (dependent and 4 independent), the 1 GiB DRAM chase (the hash's control), a point-sampled u32 texture fetch and a linearly filtered float fetch (cuTexObjectCreate loaded optionally). Each RESULT line carries the counted operations per second, lanes, registers, UTC start and end stamps for a 1 Hz power sampler, and a checksum. Gate: mingw cross-compile on igneum-build-1 under lease pool (class measure, label "ca4 research"), exit 0 with -Wall -Wextra clean, 21:20 UTC; exe sha256 49aa60c4b28ed77a54655476d0f07c06a54aa22e4f6f6ccadd4019b84cbf0268; unrun on a card (the hash lane's PC 1 queue).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs/analysis/counter-asic-4-research.md (excluded from the public export): edge = (E_card + F) / (E_mem + k F); at zero premium the stored-dataset chip keeps E_card / E_mem, 3.6x on a 5090 locked at 1,400 MHz (measured card, modelled chip), about 2x at that card's idle-plus-memory bound, 1.7x on the M5 Max; class v5 at zero shadow leaves it at 5.1x to 9.1x; the shadow buys 2.1x at k = 1 for the measured 88 W (81.8 W at the knee) and turns against us below 1.8 pJ per op; per-card classes, the refresh as a cost, proof of useful work, memory shaping and the tensor block priced out with their arithmetic; ten designs ranked with the chip edge, the 5090 and 4070 premiums, the verifier cost, agent hours and what breaks each; the recommendation in five sentences; consequences per tier.
proto-cuda/nvrtc/worker.cpp: opt-in SM-sparse variants sp<N> and sp<N>-w<W> (a persistent grid of N blocks of W warps over the dispatch's nonces; the bound kernel rewritten at compile time with exact anchors into a unit function plus a wrapper of the kernel's name; made on demand by name, never in the default race; refused on variant-5 packs and with minBlocks). Gate: mingw cross-compile on igneum-build-1 under lease pool (class measure, label "ca4 research"), exit 0 with -Wall -Wextra, 22:48 UTC; exe sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be; unrun on a card (the hash lane's PC 1 job run-ca4-pc1-ca4sparse-5090-20261007 is its first run).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>